This page walks you through the end-to-end workflow for reinforcement learning fine-tuning of Gemini models in the Google Cloud console: creating a tuning job, checking its status, retrieving the tuned-model endpoint, and running inference against it. To complete these tasks by using the Agent Platform API, see the API quick start.
Before you begin, see About reinforcement learning fine-tuning for an overview of the feature, supported models, and supported regions.
Create a reinforcement learning fine-tuning job
To create a reinforcement learning fine-tuning job by using the Google Cloud console, perform the following steps:
In the Google Cloud console, go to the Models > Tuning page.
Click Create tuned model.
In the Model details section, configure the following:
- Under Tuning method, select Reinforcement learning fine tuning.
- In the Tuned model name field, enter a name for your tuned model
(for example,
my-rl-tuned-model). - Under Base model, select Tune a foundation model.
- In the Base model drop-down list, select the model to tune (for
example,
gemini-3.5-flash). - In the Region drop-down list, select the region where the tuning job
runs (for example,
us-central1 (Iowa)). The resulting tuned model is served from theusmulti-region endpoint. For the full list of supported tuning and serving regions, see the Supported models and regions section. - Optional: Expand Advanced options to customize hyperparameters:
- Number of epochs: The number of training epochs (for example,
15). - Learning rate multiplier: The multiplier to scale the learning
rate (for example,
1.0). - Adapter size: The adapter size for parameter-efficient tuning
(for example,
16). - Samples per prompt: The number of candidate responses the model
generates per prompt during training (for example,
16). - Thinking level: The thinking level for reasoning models
(for example,
HIGH). - Batch size: The number of training examples per batch (for
example,
32). - Checkpoint interval: The frequency in steps at which
intermediate checkpoints are saved (for example,
5). - Max output token: The maximum number of output tokens generated
per sample (for example,
32768). - Evaluate interval: The frequency in steps at which evaluation
runs against the validation dataset (for example,
5).
- Number of epochs: The number of training epochs (for example,
- Click Continue.
In the Reward configuration section, configure one or more reward functions to score model responses during training:
- In the Reward name field, enter a unique name for the reward
function (for example,
my_reward_function_name). Emitted metrics in the monitoring view are prefixed with this name. - In the Reward type drop-down list, select the scorer type:
- String match reward: Evaluates generated text against references by using exact matching or regular expressions.
- LLM based reward: Uses a Gemini model as an autorater to score responses based on an evaluation prompt.
- Python function based reward: Executes custom Python code in a secure sandbox to evaluate responses.
- Fully customizable reward via Cloud Run: Calls an external HTTP endpoint hosted on Cloud Run to evaluate responses.
Configure the settings for your selected reward type. For example, for a Cloud Run reward:
- In the Cloud Run URI field, enter the HTTPS URI of your
deployed service (for example,
https://my.cloud.run.uri). - In the Parse type field, select
IDENTITY.
Alternatively, for a string match reward:
- In the Parse type field, select
REGEX_EXTRACTorIDENTITY. If you selectREGEX_EXTRACT, enter the regular expression in Regex extract expression (for example,\text{(.*)}). - In the Wrong answer reward field, enter the penalty for an
incorrect response (for example,
-1). - In the Correct answer reward field, enter the score for a
correct response (for example,
1). - Under String match expression type, select String match or JSON match.
- In the Match operation drop-down list, select the match logic (for example, Exact match).
- In the Expression field, enter the reference expression (for
example,
references.reference).
- In the Cloud Run URI field, enter the HTTPS URI of your
deployed service (for example,
In the Reward weight field, enter the relative weight for this reward function (for example,
1).Click Done.
Optional: To define a composite reward, click + Add another reward and configure additional reward functions. You can add up to 16 reward functions and assign relative weights to each.
Click Continue.
- In the Reward name field, enter a unique name for the reward
function (for example,
In the Test reward (Optional) section, validate your reward configuration before starting the tuning job:
- In the Training Example field, enter a sample JSON object containing
contentsandreferences. - In the Sample model response field, enter a candidate response, or click Generate from Training Example to generate a sample response by using the base model.
- Click Test.
- Verify that the evaluation succeeds and check the computed score. If using composite rewards, click See individual reward information to inspect the per-reward score breakdown.
- Click Continue.
- In the Training Example field, enter a sample JSON object containing
Under Tuning dataset, specify your dataset files stored in Cloud Storage:
- In the Training dataset field, enter the Cloud Storage URI
of your training dataset JSONL file (for example,
gs://path/to/my/training_dataset.jsonl). - In the Validation dataset field, enter the
Cloud Storage URI of your validation dataset JSONL file
(for example,
gs://path/to/my/eval_dataset.jsonl). - Click Start tuning.
- In the Training dataset field, enter the Cloud Storage URI
of your training dataset JSONL file (for example,
Check the reinforcement learning fine-tuning job status
After you click Start tuning, the console redirects to the Models > Tuning page. The tuning jobs table displays your newly created job near the top of the list with the following details:
- Model name: The display name you specified for the tuned model.
- Status: The current state of the tuning job (such as Pending, Running, or Succeeded).
- Method:
Reinforcement Learning. - Base model: The foundation or pre-tuned model.
- Region: The location where the tuning job is running.
Click the model name to open the tuning job details page. The page provides the following tabs:
- Monitor: Surfaces tuning progress and interactive metrics charts:
- Tuning progress: Shows whether the job is running or completed.
- Combined charts: When a validation dataset is provided, training and
validation curves (such as
/train_mean_rewardand/eval_mean_reward) are plotted together on the same charts. - Collapsible chart groups: Metric charts are organized into groups,
including general tuning metrics (mean reward, generation token length,
thinking token length, sampling latency, and reward latency) and
per-reward metrics prefixed by
${reward_name}. - Checkpoint annotations: Intermediate checkpoints saved during tuning are annotated directly on the charts along the training timeline.
- Filter bar: Use the filter bar to search metrics by name or filter
by category (
Show Category). When filtered, the filter bar remains sticky at the top of the view. - Checkpoints table: Lists saved intermediate checkpoints with step numbers and metric values. Use the column selector to choose which columns to display.
- Dataset: Displays dataset details, sample conversations, and reference values.
- Details: Summarizes the job configuration, including the base model, tuning method, hyperparameters, and reward configurations. Click View details on a reward configuration card to view non-default parameter values.
For the full list of emitted metrics and how to interpret them, see the Metrics and monitoring page.
Training time
Training time is affected by the following factors:
- Training and validation dataset size: For details, see Tuning dataset for reinforcement learning fine-tuning.
- Hyperparameters: Includes
samplesPerPrompt, batch size, epoch count, and learning rate multiplier. For details, see Hyperparameters for reinforcement learning fine-tuning.
Depending on your setup, a Gemini reinforcement learning fine-tuning job can run for hours to days.
Get the tuned-model endpoint
After the tuning job reaches Succeeded, the model checkpoint is deployed to an endpoint. You can view and manage the deployed tuned model directly in the Google Cloud console:
- In the header of the tuning job details page, click View model details. The Model Registry page opens with details for the tuned model.
- On the model details page, find the Deploy & test tab to view the
deployed endpoint resource path (in the format
projects/{PROJECT_ID}/locations/{LOCATION}/endpoints/{ENDPOINT_ID}). - Alternatively, on the Monitor tab of the tuning job details page, locate the Checkpoints table to inspect intermediate checkpoints. You can deploy and evaluate individual checkpoints independently.
If the tuning job ran in us-central1, the tuned model is served from the
us multi-region endpoint.
Run inference on the tuned model
You can test and run inference on the tuned model directly in the Google Cloud console, do the following:
In the header of the tuning job details page, click Test.
Agent Studio opens with your tuned model pre-selected.
In the prompt field, enter an input prompt (for example,
"Why is the sky blue?"or a question from your target domain).Click Run (or Send) to generate a response.
Review the generated output to evaluate how well the model adheres to your desired reasoning style, formatting, and reward objectives.
What's next
- Learn more about reinforcement learning fine-tuning jobs.
- Explore supported reward functions.
- Configure hyperparameters for reinforcement learning.
- Learn how to track metrics and monitoring.
- Try the API quick start to create tuning jobs programmatically.