Cloud Run lets you attach GPUs to deploy open models and serve AI inference. You can train, fine-tune, and run models with accelerated performance.
If you are new to AI concepts, see GPUs for AI. To learn more about configuration options, see GPU support for services, jobs, and worker pools.
Tutorials for services
- Run Gemma on Cloud Run
- Run LLM inference on Cloud Run GPUs with Gemma and Ollama
- Run OpenCV on Cloud Run with GPU acceleration
- Run LLM inference on Cloud Run GPUs with Hugging Face Transformers.js
Tutorials for jobs
- Fine tune LLMs using GPUs with Cloud Run jobs
- Run batch inference using GPUs on Cloud Run jobs
- GPU-accelerated video transcoding with FFmpeg on Cloud Run jobs