By default, Cloud Run optimizes for high performance with a utilization target of 60% for both CPU and concurrency, and scales the number of instances automatically to handle all incoming requests. However, for some use cases, you might want the ability to configure which scaling factors to use, such as CPU only, and set custom targets for utilization.
Cloud Run provides scaling controls to give you more ownership of your service's scaling behaviors, letting you make informed decisions about scaling your workload according to your requirements. You can configure the following custom utilization targets:
- Target utilization for CPU-based scaling
- Target utilization for concurrency-based scaling
With scaling controls, you can optimize costs and improve predictability for your services. For more information about the default autoscaling behavior of Cloud Run services, see About instance autoscaling in Cloud Run services.
Configuration limits
The following limits apply to custom scaling targets:
| Scaling driver | Default % | Minimum configurable % | Maximum configurable % |
|---|---|---|---|
CPU target utilization |
60% | 10% | 90% |
Concurrency target utilization |
60% | 10% | 95% |
Configure custom targets
Define custom utilization targets to optimize costs or improve performance for your workloads by configuring specific CPU and concurrency utilization targets within the configuration limits.
Any configuration change leads to the creation of a new revision. Subsequent revisions will also automatically get this configuration setting unless you make explicit updates to change it.
Even if you configure custom concurrency targets or disable CPU-based scaling, Adaptive Concurrency Tuning (ACT) remains active. For details, see About instance autoscaling.
You can configure scaling controls using the Google Cloud console, gcloud CLI, YAML, or Terraform when you deploy a new revision.
Console
In the Google Cloud console, go to the Cloud Run Services page:
If you are configuring a new service, click Deploy container to display the Create service page.
If you are configuring an existing service, click the service to open its Service details page, then click the Scaling tab.
Locate the Service scaling section. Make sure Auto scaling is selected. Expand the Customize autoscaling factors section to configure the following utilization targets:
To configure target CPU utilization, click CPU utilization and enter a value from
10to90.To configure target concurrency utilization, click Concurrent request utilization and enter a value from
10to95.Click Done under each utilization configuration.
Click Create for a new service. Click View diff & redeploy, then Deploy changes for an existing service.
gcloud
Update the target CPU utilization and the target concurrency utilization values of a given revision by
running the gcloud run services update command.
To update the target CPU utilization, run the following command:
gcloud run services update SERVICE --scaling-cpu-target=CPU_TARGET
Replace the following:
SERVICE: the name of your service.
CPU_TARGET: the target for CPU utilization. Specify a value from
0.1to0.90. You can only configure up to two digits after the decimal point.
To update the target concurrency utilization, run the following command:
gcloud run services update SERVICE --scaling-concurrency-target=CONCURRENCY_TARGET
Replace the following:
SERVICE: the name of your service.
CONCURRENCY_TARGET: the target for concurrency utilization. Specify a value from `0.1` to `0.95`. You can only configure up to two digits after the decimal point.
To update both the target CPU and the concurrency utilization, run the following command:
gcloud run services update SERVICE --scaling-cpu-target=CPU_TARGET \ --scaling-concurrency-target=CONCURRENCY_TARGET
Replace the following:
- SERVICE: the name of your service.
- CPU_TARGET: the target for CPU utilization. Specify a value
from
0.1to0.90. You can only configure up to two digits after the decimal point. - CONCURRENCY_TARGET: the target for concurrency utilization. Specify a value from `0.1` to `0.95`. You can only configure up to two digits after the decimal point.
YAML
If you are creating a new service, skip this step. If you are updating an existing service, download its YAML configuration:
gcloud run services describe SERVICE --format export > service.yaml
To update the target CPU and concurrency utilization, add the
run.googleapis.com/scaling-cpu-targetandrun.googleapis.com/scaling-concurrency-targetattributes:apiVersion: serving.knative.dev/v1 kind: Service metadata: name: SERVICE spec: template: metadata: annotations: run.googleapis.com/scaling-cpu-target: 'CPU_TARGET' run.googleapis.com/scaling-concurrency-target: 'CONCURRENCY_TARGET'
Replace the following:
- SERVICE: the name of your service.
- CPU_TARGET: the target for CPU utilization. Specify a value
from
0.1to0.90. You can only configure up to two digits after the decimal point. - CONCURRENCY_TARGET: the target for concurrency utilization. Specify a value from `0.1` to `0.95`. You can only configure up to two digits after the decimal point.
Create or update the service using the following command:
gcloud run services replace service.yaml
The
gcloud run services replacecommand defaults to usingservice.yamlfile if present.
Terraform
To learn how to apply or remove a Terraform configuration, see Basic Terraform commands.
Add the following to agoogle_cloud_run_v2_service
resource in your Terraform configuration:resource "google_cloud_run_v2_service" "default" {
name = "SERVICE"
location = "REGION"
template {
scaling {
cpu_utilization = CPU_TARGET
concurrency_utilization = CONCURRENCY_TARGET
}
containers {
image = "IMAGE_URL"
}
}
}
Replace the following:
- SERVICE: the name of your service.
- REGION: the Google Cloud region, for example,
europe-west1. - CPU_TARGET: the target for CPU utilization. Specify a value
from
0.1to0.90. You can only configure up to two digits after the decimal point. - CONCURRENCY_TARGET: the target for concurrency utilization. Specify a value from `0.1` to `0.95`. You can only configure up to two digits after the decimal point.
IMAGE_URL: a reference to the container image, for example,us-docker.pkg.dev/cloudrun/container/hello:latest. If you use Artifact Registry, the repository REPO_NAME must already be created. The URL follows the format ofLOCATION-docker.pkg.dev/PROJECT_ID/REPO_NAME/PATH:TAG
Disable scaling controls
You can disable either CPU utilization or concurrency utilization targets, but not both. One scaling driver must always be active. To opt out of scaling controls, restore the default utilization values instead of disabling them. When you disable a scaling driver, Cloud Run ignores that metric while making scaling decisions.
You can disable scaling controls using the Google Cloud console, gcloud CLI, YAML, or Terraform when you deploy a new revision.
Console
In the Google Cloud console, go to the Cloud Run Services page:
Click the service to open its Service details page, then click the Scaling tab.
Locate the Service scaling section. Make sure Auto scaling is selected. Expand the Customize autoscaling factors section:
To scale only by CPU, click the delete icon next to Concurrent request utilization, and enter a value for target CPU utilization if it's missing.
To scale only by concurrency, click the delete icon next to CPU utilization, and enter a value for target concurrency utilization if it's missing.
Click Done under each utilization configuration.
Click Create for a new service. Click View diff & redeploy, then Deploy changes for an existing service.
gcloud
You can disable either the target CPU utilization or the target concurrency utilization by
running the gcloud run services update command.
To scale only by CPU, disable the concurrency target by running the following command:
gcloud run services update SERVICE --scaling-concurrency-target=disabled
Replace SERVICE with the name of your service.
To scale only by concurrency, disable the CPU target by running the following command:
gcloud run services update SERVICE --scaling-cpu-target=disabled
Replace SERVICE with the name of your service.
YAML
If you are creating a new service, skip this step. If you are updating an existing service, download its YAML configuration:
gcloud run services describe SERVICE --format export > service.yaml
To scale only by CPU, disable the concurrency target by setting the
run.googleapis.com/scaling-concurrency-targetattribute todisabled:apiVersion: serving.knative.dev/v1 kind: Service metadata: name: SERVICE spec: template: metadata: annotations: run.googleapis.com/scaling-concurrency-target: disabled
Replace SERVICE with the name of your service.
To scale only by concurrency, disable the CPU target by setting the
run.googleapis.com/scaling-cpu-targetattribute todisabled:apiVersion: serving.knative.dev/v1 kind: Service metadata: name: SERVICE spec: template: metadata: annotations: run.googleapis.com/scaling-cpu-target: disabled
Replace SERVICE with the name of your service.
Create or update the service using the following command:
gcloud run services replace service.yaml
The
gcloud run services replacecommand defaults to usingservice.yamlfile if present.
Terraform
To learn how to apply or remove a Terraform configuration, see Basic Terraform commands.
Add the following to agoogle_cloud_run_v2_service
resource in your Terraform configuration:To scale only by CPU, disable the concurrency target by setting
concurrency_utilizationto0:resource "google_cloud_run_v2_service" "default" { name = "SERVICE" location = "REGION" template { scaling { cpu_utilization = CPU_TARGET concurrency_utilization = 0 } containers { image = "IMAGE_URL" } } }Replace the following:
- SERVICE: the name of your service.
- REGION: the Google Cloud region, for example,
europe-west1. - CPU_TARGET: the target for CPU utilization. Specify a value
from
0.1to0.90. You can only configure up to two digits after the decimal point. IMAGE_URL: a reference to the container image, for example,us-docker.pkg.dev/cloudrun/container/hello:latest. If you use Artifact Registry, the repository REPO_NAME must already be created. The URL follows the format ofLOCATION-docker.pkg.dev/PROJECT_ID/REPO_NAME/PATH:TAG
To scale only by concurrency, disable the CPU target by setting
cpu_utilizationto0:resource "google_cloud_run_v2_service" "default" { name = "SERVICE" location = "REGION" template { scaling { cpu_utilization = 0 concurrency_utilization = CONCURRENCY_TARGET } containers { image = "IMAGE_URL" } } }Replace the following:
- SERVICE: the name of your service.
- REGION: the Google Cloud region, for example,
europe-west1. - CONCURRENCY_TARGET: the target for concurrency utilization. Specify a value from `0.1` to `0.95`. You can only configure up to two digits after the decimal point.
IMAGE_URL: a reference to the container image, for example,us-docker.pkg.dev/cloudrun/container/hello:latest. If you use Artifact Registry, the repository REPO_NAME must already be created. The URL follows the format ofLOCATION-docker.pkg.dev/PROJECT_ID/REPO_NAME/PATH:TAG
Restore to default values
When you restore the target CPU or the target concurrency utilization values to default, Cloud Run uses the default utilization target of 60% instead of your custom targets. You can restore scaling controls to default using the Google Cloud console, gcloud CLI, YAML, or Terraform when you deploy a new revision.
Console
In the Google Cloud console, go to the Cloud Run Services page:
Click the service to open its Service details page, then click the Scaling tab.
Locate the Service scaling section. Make sure Auto scaling is selected. Expand the Customize autoscaling factors section to configure the following utilization targets:
If the CPU utilization and Concurrent request utilization targets were removed, click Add a signal to add each back.
Set the values for CPU utilization and Concurrent request utilization to
60.Click Done under each utilization configuration.
Click Create for a new service. Click View diff & redeploy, then Deploy changes for an existing service.
gcloud
Restore the target CPU utilization and the target concurrency utilization to their defaults by
running the gcloud run services update command.
To restore the target CPU utilization to its default value, run the following command:
gcloud run services update SERVICE --scaling-cpu-target=default
Replace SERVICE with the name of your service.
To restore the target concurrency utilization to its default value, run the following command:
gcloud run services update SERVICE --scaling-concurrency-target=default
Replace SERVICE with the name of your service.
To restore both the target CPU utilization and the target concurrency to their default values, run the following command:
gcloud run services update SERVICE --scaling-cpu-target=default \ --scaling-concurrency-target=default
Replace SERVICE with the name of your service.
YAML
If you are creating a new service, skip this step. If you are updating an existing service, download its YAML configuration:
gcloud run services describe SERVICE --format export > service.yaml
To restore CPU and concurrency utilization to their default targets, remove the
run.googleapis.com/scaling-cpu-targetandrun.googleapis.com/scaling-concurrency-targetattributes from your YAML file:apiVersion: serving.knative.dev/v1 kind: Service metadata: name: SERVICE spec: template: metadata: # Remove the scaling target annotations to restore defaults ...
Replace SERVICE with the name of your service.
Create or update the service using the following command:
gcloud run services replace service.yaml
The
gcloud run services replacecommand defaults to usingservice.yamlfile if present.
Terraform
To learn how to apply or remove a Terraform configuration, see Basic Terraform commands.
Add the following to agoogle_cloud_run_v2_service
resource in your Terraform configuration:To restore CPU and concurrency utilization to their default targets, remove
the cpu_utilization and concurrency_utilization attributes from the scaling
block in your Terraform configuration:
resource "google_cloud_run_v2_service" "default" {
name = "SERVICE"
location = "REGION"
template {
scaling {
# Remove the scaling target attributes to restore defaults
}
containers {
image = "IMAGE_URL"
}
}
}
Replace the following:
- SERVICE: the name of your service.
- REGION: the Google Cloud region, for example,
europe-west1. IMAGE_URL: a reference to the container image, for example,us-docker.pkg.dev/cloudrun/container/hello:latest. If you use Artifact Registry, the repository REPO_NAME must already be created. The URL follows the format ofLOCATION-docker.pkg.dev/PROJECT_ID/REPO_NAME/PATH:TAG
View scaling configuration
You can view your scaling configuration using the Google Cloud console or gcloud CLI.
Console
In the Google Cloud console, go to the Cloud Run Services page:
Click your service to open the Service details panel.
Click the Scaling tab to view scaling settings.
gcloud
Use the following command:
gcloud run services describe SERVICE
Replace SERVICE with the name of your service.
Locate the value for the Target CPU utilization: and Target concurrency utilization: settings in the returned configuration.
Best practices
You can optimize costs and prevent overscaling by decreasing the number of instances, or you can improve performance by scaling more aggressively in response to specific drivers. To determine the optimal utilization targets for your workload, use the following strategies:
Before adjusting targets, identify which metric is triggering your service to scale. Follow these steps to identify the scaling metric:
Go to Metrics Explorer in the Google Cloud console to review the monitoring chart for your scaling driver.
Search and select the
run.googleapis.com/scaling/recommended_instancesmetric, and set Aggregation to Unaggregated to view the metric grouped by scaling driver.
The driver with the highest value is the one controlling your service's instance count. If you want a different driver to take priority, or if you want to scale more or less aggressively, adjust the utilization target for that specific driver.
If Adaptive Concurrency Tuning (ACT) is the scaling driver, it indicates that 1-second CPU utilization is exceeding 90% for individual instances and Cloud Run is dynamically limiting request concurrency to protect your service. To mitigate ACT driving your scaling, consider increasing the CPU allocation or adjusting your concurrency settings. For more information, see About instance autoscaling.
Adjust targets incrementally, and wait for a few minutes between adjustments to observe the effect on performance.
Use traffic splitting to test new scaling targets by directing a small percentage of your traffic to a separate revision before rolling them out to your entire service.
About low utilization targets
Lowering your utilization target to the minimum of 0.1 (10%) significantly
changes how your service scales.
Benefits of setting a low utilization target include:
High service availability: Your service scales up much earlier, maintaining a large buffer of idle capacity to handle sudden traffic spikes without latency hits.
Faster scaling at low instance counts: Services scale more reliably before hitting high-utilization bottlenecks.
Drawbacks of setting low utilization targets include:
- Potential for cost increase: You run more instances than strictly necessary for your current load, leading to higher billing.
- More frequent scaling decisions: At lower utilizations, Cloud Run has a lower tolerance, and doesn't wait as long before scaling.
What's next
- To learn more about autoscaling concepts and behaviors, see About instance autoscaling in Cloud Run services.
- To learn about other scaling options, see manual scaling.
- To manage the maximum number of instances of your Cloud Run services, see Setting a maximum number of instances.
- To manage the maximum number of simultaneous requests handled by each instance, see Setting concurrency.
- To optimize your concurrency setting, see development tips for tuning concurrency.
- To specify an idle instance to keep running to minimize latency or cold starts
on first requests, see
Using
min-instanceto enable idle instances.