This page provides best practices for optimizing performance when using a Cloud Run service with GPU for AI inference, focusing on large language models (LLMs). To build and deploy a Cloud Run service that can respond in real time to scaling events, you should:

- Use models that load fast and require minimal transformation into GPU-ready structures, and optimize how they are loaded.
- Use configurations that allow for maximum, efficient, concurrent execution to reduce the number of GPUs needed to serve a target request per second while keeping costs down.

## Recommended ways to load large ML models on Cloud Run

Google recommends [downloading ML models from Cloud Storage](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-storage) and accessing them through the Google Cloud CLI. You might alternatively store models inside container images, but this method is best suited for smaller models less than 10 GB.

### Storing and loading ML models trade-offs

Here is a comparison of the options:

|---|---|---|---|---|
| **Model location** | **Deploy time** | **Development experience** | **Container startup time** | **Storage cost** |
| [Cloud Storage](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-storage), downloaded concurrently using the Google Cloud CLI command [`gcloud storage cp`](https://docs.cloud.google.com/sdk/gcloud/reference/storage/cp) or the Cloud Storage API as shown in the [transfer manager concurrent download code sample](https://docs.cloud.google.com/storage/docs/samples/storage-transfer-manager-download-chunks-concurrently). | Fastest. Model downloaded during container startup. Ensure the Cloud Run instance has sufficient RAM allocated to store the model files. | Slightly more difficult to set up, because you'll need to either install the Google Cloud CLI on the image or update your code to use the Cloud Storage API. For more information on how to fetch credentials from the metadata server, see [Introduction to service identity](https://docs.cloud.google.com/run/docs/securing/service-identity#access-tokens). | Fast when you use [network optimizations](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-storage). The Google Cloud CLI downloads the model file in parallel. | One copy in Cloud Storage. |
| [Cloud Storage](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-storage), loaded using Cloud Storage FUSE volume mount | Faster. Model downloaded during container startup. | Not difficult to set up, does not require changes to the docker image. | Fast when you use [network optimizations](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-storage). | One copy in Cloud Storage. |
| [Container image](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-container) | Fast. An image containing a large model will take longer to import into Cloud Run. | You'll need to build a new image every time you want to use a different model. Changes to the container image will require redeployment, which may be slow for large images. | Depends on the size of the model. For very large models, use Cloud Storage for more predictable but slower performance. | Potentially multiple copies in Artifact Registry. |
| [Internet](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#model-internet) | Slow. Model downloaded during container startup. | Typically simpler (many frameworks download models from central repositories). | Typically poor and unpredictable: - Frameworks may apply model transformations during initialization. (You should do this at build time). - Model host and libraries for downloading the model may not be efficient. - There is reliability risk associated with downloading from the internet. Your service could fail to start if the download target is down, and the underlying model downloaded could change, which decreases quality. We recommend hosting in your own Cloud Storage bucket. | Depends on the model hosting provider. |

### Store models in Cloud Storage

To optimize ML model loading when loading ML models from Cloud Storage,
either using [Cloud Storage volume mounts](https://docs.cloud.google.com/run/docs/configuring/services/cloud-storage-volume-mounts)
or directly using the Cloud Storage API or command line, you must use [Direct VPC](https://docs.cloud.google.com/run/docs/configuring/vpc-direct-vpc) with
the egress setting value set to `all-traffic`, along with [Private Google Access](https://docs.cloud.google.com/vpc/docs/configure-private-google-access).

For an additional cost, using [Rapid Cache](https://docs.cloud.google.com/storage/docs/rapid/rapid-cache) can reduce model loading latency by efficiently caching data on SSDs for faster reads.

To reduce model read times, try the following mount options to enable Cloud Storage FUSE features:

- `cache-dir`: Enable the [file caching feature](https://docs.cloud.google.com/storage/docs/cloud-storage-fuse/file-caching) with an [in-memory volume mount](https://docs.cloud.google.com/run/docs/configuring/services/in-memory-volume-mounts) to use as the underlying directory to persist files. Set the `cache-dir` mount option value to the in-memory volume name in the format, `cr-volume:{volume name}`. For example, if you have an in-memory volume named `in-memory-1` that you want to use as the cache directory, specify `cr-volume:in-memory-1`. When this value is set, you can also set other `file-cache` [flags available to configure for cache](https://docs.cloud.google.com/run/docs/configuring/services/cloud-storage-volume-mounts#configure-cache).
- `enable-buffered-read`: Set the [`enable-buffered-read`](https://docs.cloud.google.com/storage/docs/cloud-storage-fuse/cli-options#enable-buffered-read) field to `true` for asynchronous prefetching of parts of a Cloud Storage object into an in-memory buffer. This allows subsequent reads to be served from the buffer instead of requiring network calls. When you configure this field, you can also set the [`read-global-max-blocks`](https://docs.cloud.google.com/storage/docs/cloud-storage-fuse/cli-options#read-global-max-blocks) field to configure the maximum number of blocks available for buffered reads across all file handles.

When both `cache-dir` and `enable-buffered-read` are used, `cache-dir` will take precedence. Note that enabling any of these features will change the resource accounting of the Cloud Storage FUSE process to be counted under container memory limits. Consider raising the container memory limit by following [instructions on how to configure memory limits](https://docs.cloud.google.com/run/docs/configuring/services/memory-limits#configure-services).

### Store models in container images

By storing the ML model in the container image, model loading will benefit from Cloud Run's optimized container streaming infrastructure.
However, building container images that include ML models is a resource-intensive process,
especially when working with large models. In particular, the build process can become
bottlenecked on network throughput. When using Cloud Build, we recommend
using a more powerful build machine with increased compute and networking
performance. To do this, build an image using a
[build config file](https://docs.cloud.google.com/build/docs/build-push-docker-image#build_an_image_using_a_build_config_file)
that has the following steps:

```yaml
steps:
- name: 'gcr.io/cloud-builders/docker'
  args: ['build', '-t', 'IMAGE', '.']
- name: 'gcr.io/cloud-builders/docker'
  args: ['push', 'IMAGE']
images:
- IMAGE
options:
 machineType: 'E2_HIGHCPU_32'
 diskSizeGb: '500'
 
```

You can create one model copy per image if the layer containing the model is
distinct between images (different hash). There could be additional Artifact Registry
cost because there could be one copy of the model per image if your model layer
is unique across each image.

### Load models from the internet

To optimize ML model loading from the internet, [route all traffic through
the vpc network](https://docs.cloud.google.com/run/docs/configuring/vpc-direct-vpc) with the egress setting
value set to `all-traffic` and set up [Cloud NAT](https://docs.cloud.google.com/nat/docs/overview) to reach the public internet at high bandwidth.

## Build, deployment, runtime, and system design considerations

The following sections describe considerations for build, deploy, runtime and system design.

### At build time

The following list shows considerations you need to take into account when you
are planning your build:

- Choose a good base image. You should start with an image from the [Deep Learning Containers](https://docs.cloud.google.com/deep-learning-containers/docs/choosing-container) or the [NVIDIA container registry](https://catalog.ngc.nvidia.com/containers) for the ML framework you're using. These images have the latest performance-related packages installed. We don't recommend creating a custom image.
- Choose 4-bit quantized models to maximize concurrency unless you can prove they affect result quality. Quantization produces smaller and faster models, reducing the amount of GPU memory needed to serve the model, and can increase parallelism at run time. Ideally, the models should be trained at the target bit depth rather than quantized down to it.
- Pick a model format with fast load times to minimize container startup time, such as GGUF. These formats more accurately reflect the target quantization type and require less transformations when loaded onto the GPU. For security reasons, don't use pickle-format checkpoints.
- Create and warm LLM caches at build time. Start the LLM on the build machine while building the docker image. Enable prompt caching and feed common or example prompts to help warm the cache for real-world use. Save the outputs it generates to be loaded at runtime.
- Save your own inference model that you generate during build time. This saves significant time compared to loading less efficiently stored models and applying transforms like quantization at container startup.

### At deployment

The following list shows considerations you need to take into account when you
are planning your deployment:

1. Make sure you [set service concurrency accurately](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#max-concurrent-requests) in Cloud Run.
2. Adjust your [startup probes](https://docs.cloud.google.com/run/docs/configuring/healthchecks) based on your configuration.

<br />

Startup probes determine whether the container has started and is ready to accept
traffic. Consider these key points when configuring startup probes:

- Adequate startup time: Allow sufficient time for your container, including models, to fully initialize and load.
- Model readiness verification: Configure your probe to pass only when your application is ready to serve requests. Many serving engines automatically achieve this when the model is loaded into GPU memory, preventing premature requests.

Note that Ollama can open a TCP port before a model is loaded. To address this:

- Preload models: Refer to the [Ollama documentation](https://github.com/ollama/ollama/blob/main/docs/faq.md#how-can-i-preload-a-model-into-ollama-to-get-faster-response-times) for guidance on preloading your model during startup.

### At run time

- Actively manage your supported context length. The smaller the context window you support, the more queries you can support running in parallel. The details of how to do this depend on the framework.
- Use the LLM caches you generated at build time. Supply the same flags you used during build time when you generated the prompt and prefix cache.
- Load from the saved model you just wrote. See [Storing and loading models trade-offs](https://docs.cloud.google.com/run/docs/configuring/services/gpu-best-practices#loading-storing-models-tradeoff) for a comparison on how to load the model.
- Consider using a quantized key-value cache if your framework supports it. This can reduce per-query memory requirements and allows for configuration of more parallelism. However, it can also impact quality.
- Tune the amount of GPU memory to reserve for model weights, activations and key-value caches. Set it as high as you can without getting an out-of-memory error.
- Check to see whether your framework has any options for improving container startup performance (for example, using model loading parallelization).
- Configure your concurrency correctly inside your service code. Make sure your service code is configured to work with your Cloud Run service concurrency settings.

### At the system design level

- Add semantic caches where appropriate. In some cases, caching whole queries and responses can be a great way of limiting the cost of common queries.
- Control variance in your preambles. Prompt caches are only useful when they contain the prompts in sequence. Caches are effectively prefix-cached. Insertions or edits in the sequence mean that they're either not cached or only partially present.

## Autoscaling and GPUs

If you use the default Cloud Run autoscaling, [Cloud Run automatically scales the number of instances](https://docs.cloud.google.com/run/docs/about-instance-autoscaling) of each revision based on factors such as CPU
utilization and request concurrency. However, Cloud Run
does *not* automatically scale the number of instances based on GPU utilization.

For a revision with a GPU, if the revision does not have significant CPU usage,
Cloud Run scales out for request concurrency. To achieve optimal
scaling for request concurrency, you must set an optimal
[maximum concurrent requests per instance](https://docs.cloud.google.com/run/docs/about-concurrency), as
described in the next section.

### Maximum concurrent requests per instance

The [maximum concurrent requests per instance](https://docs.cloud.google.com/run/docs/configuring/concurrency)
setting controls the maximum number of requests Cloud Run sends to a
single instance at once. You must [tune concurrency](https://docs.cloud.google.com/run/docs/tips/general#tuning-concurrency)
to match the maximum concurrency the code inside each instance can handle with
good performance.

#### Maximum concurrency and AI workloads

When running an AI inference workload on a GPU in each instance, the maximum
concurrency that the code can handle with good performance depends on specific
framework and implementation details. The following impacts how you set the
optimal maximum concurrent requests setting:

- Number of model instances loaded onto the GPU
- Number of parallel queries per model
- Use of batching
- Specific batch configuration parameters
- Amount of non-GPU work

If maximum concurrent requests is set too high, requests might end up waiting
inside the instance for access to the GPU, which leads to increased latency. If
maximum concurrent requests is set too low, the GPU might be underutilized
causing Cloud Run to scale out more instances than necessary.

A rule of thumb for configuring maximum concurrent requests for AI
workloads is:

`(Number of model instances * parallel queries per model) + (number of model instances * ideal batch size)`

For example, suppose an instance loads `3` model instances onto the GPU, and
each model instance can handle `4` parallel queries. The ideal batch size is
also `4` because that is the number of parallel queries each model instance
can handle. Using the rule of thumb, you would set maximum concurrent requests
`24`: (`3` \* `4`) + (`3` \* `4`).

Note that this formula is just a rule of thumb. The ideal maximum concurrent
requests setting depends on the specific details of your
implementation. To achieve your actual optimal performance, we recommend
[load testing](https://docs.cloud.google.com/run/docs/about-load-testing) your service with different maximum
concurrent requests settings to evaluate which option performs best.

### Throughput versus latency versus cost tradeoffs

See [Throughput versus latency versus costs tradeoffs](https://docs.cloud.google.com/run/docs/tips/general#throughput-latency-cost-tradeoff)
for the impact of maximum concurrent requests on throughput, latency, and cost.
Note that all Cloud Run services using GPUs must have
[instance-based billing](https://docs.cloud.google.com/run/docs/configuring/billing-settings) configured.