This page provides a full list of managed rubric-based metrics offered by the Gen AI evaluation service, which you can use in the GenAI Client in Agent Platform SDK.
For more information about test-driven evaluation, see Define your evaluation metrics.
Overview
The Gen AI evaluation service offers a list of managed rubric-based metrics for the test-driven evaluation framework:
For metrics with adaptive rubrics, most of them include both the workflow for rubric generation for each prompt and rubric validation. You can run them separately if needed. See Run an evaluation for details.
For metrics with static rubrics, no per-prompt rubrics are generated. For details regarding the intended outputs, see Metric details.
Each managed rubric-based metric has a versioning number. The metric uses the latest version by default, but you can pin to a specific version if needed:
from vertexai import types
text_quality_metric = types.RubricMetric.TEXT_QUALITY
general_quality_v1 = types.RubricMetric.GENERAL_QUALITY(version='v1')
Model versions and regional availability
The latest versions of most managed metrics (for example,
general_quality_v2, instruction_following_v2, safety_v3, and
final_response_match_v3) use Gemini 3.5 Flash as the judge model.
Agent metrics that generate rubrics, such as final_response_quality_v2 and
tool_use_quality_v2, also use Gemini 3.1 Pro for rubric generation.
The previous versions use Gemini 2.5 Flash and
Gemini 2.5 Pro. Each metric's previous version and its LLM calls are
listed in Metric details.
The latest versions are only available in locations where Gemini 3.5 Flash is available. If you run evaluations in a location where Gemini 3.5 Flash isn't available, pin the metric to its previous version, as shown in the preceding example.
Backward compatibility
For metrics offered as a Metric prompt templates, you can still access the pointwise metrics through the GenAI Client in Agent Platform SDK through the same approach. Pairwise metrics are not supported by the GenAI Client in Agent Platform SDK, but see Run an evaluation to compare two models in the same evaluation.
from vertexai import types
# Access metrics represented by metric prompt template examples
coherence = types.RubricMetric.COHERENCE
fluency = types.RubricMetric.FLUENCY
Managed metrics details
This section lists managed metrics with details such as their type, required inputs, and expected output:
- General quality
- Text quality
- Instruction following
- Grounding
- Safety
- Multi-turn general quality
- Multi-turn text quality
- Agent final response match
- Agent final response reference free
- Agent final response quality
- Agent hallucination
- Agent tool use quality
- Agent multi-turn task success
- Agent multi-turn tool use quality
- Agent multi-turn trajectory quality
- Gecko text-to-image quality
- Gecko text-to-video quality
General quality
| Latest version | general_quality_v2 |
| Type | Adaptive rubrics |
| Description | A comprehensive adaptive rubrics metric that evaluates the overall quality of a model's response. It automatically generates and assesses a broad range of criteria based on the prompt's content. This is the recommended starting point for most evaluations. |
| How to access in SDK | types.RubricMetric.GENERAL_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 6 calls to Gemini 3.5 Flash |
| Previous version | general_quality_v1: 6 calls to Gemini 2.5 Flash |
Text quality
| Latest version | text_quality_v2 |
| Type | Adaptive rubrics |
| Description | A targeted adaptive rubrics metric that specifically evaluates the linguistic quality of the response. It assesses aspects like fluency, coherence, and grammar. |
| How to access in SDK | types.RubricMetric.TEXT_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 6 calls to Gemini 3.5 Flash |
| Previous version | text_quality_v1: 6 calls to Gemini 2.5 Flash |
Instruction following
| Latest version | instruction_following_v2 |
| Type | Adaptive rubrics |
| Description | A targeted adaptive rubrics metric that measures how well the response adheres to the specific constraints and instructions given in the prompt. |
| How to access in SDK | types.RubricMetric.INSTRUCTION_FOLLOWING |
| Input |
|
| Output |
|
| Number of LLM calls | 6 calls to Gemini 3.5 Flash |
| Previous version | instruction_following_v1: 6 calls to Gemini 2.5 Flash |
Grounding
| Latest version | grounding_v2 |
| Type | Static rubrics |
| Description | A score-based metric that checks for factuality and consistency. It verifies that the model's response is grounded based on the context. |
| How to access in SDK | types.RubricMetric.GROUNDING |
| Input |
|
| Output |
0-1. If any sentence is labeled unsupported or contradictory, the score is 0. Otherwise, the score represents the ratio of sentences labeled supported or no_rad to the total number of sentences.
The explanation field is a JSON string containing a list of per-sentence objects with the following schema:
[ { "sentence": "string", "label": "supported | unsupported | contradictory | no_rad", "rationale": "string", "excerpt": "string or null" } ]
|
| Number of LLM calls | 1 call to Gemini 3.5 Flash |
| Previous version | grounding_v1: 1 call to Gemini 2.5 Flash |
Safety
| Latest version | safety_v3 |
| Type | Static rubrics |
| Description |
A score-based metric that assesses whether the model's response violated one or more safety policies.
safety_v3 checks the following policies:
safety_v1 checks the following policies:
|
| How to access in SDK | types.RubricMetric.SAFETY |
| Input |
|
| Output |
|
| Number of LLM calls | 31 calls to Gemini 3.5 Flash |
| Previous version | safety_v1: 10 calls to Gemini 2.5 Flash |
Multi-turn general quality
| Latest version | multi_turn_general_quality_v2 |
| Type | Adaptive rubrics |
| Description | An adaptive rubrics metric that evaluates the overall quality of a model's response within the context of a multi-turn dialogue. |
| How to access in SDK | types.RubricMetric.MULTI_TURN_GENERAL_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 6 calls to Gemini 3.5 Flash |
| Previous version | multi_turn_general_quality_v1: 6 calls to Gemini 2.5 Flash |
Multi-turn text quality
| Latest version | multi_turn_text_quality_v2 |
| Type | Adaptive rubrics |
| Description | An adaptive rubrics metric that evaluates the text quality of a model's response within the context of a multi-turn dialogue. |
| How to access in SDK | types.RubricMetric.TEXT_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 6 calls to Gemini 3.5 Flash |
| Previous version | multi_turn_text_quality_v1: 6 calls to Gemini 2.5 Flash |
Agent final response match
| Latest version | final_response_match_v3 |
| Type | Static rubrics |
| Description | A metric that evaluates the quality of an AI agent's final answer by comparing it to a provided reference answer (ground truth). |
| How to access in SDK | types.RubricMetric.FINAL_RESPONSE_MATCH |
| Input |
|
| Output |
Score
|
| Number of LLM calls | 5 calls to Gemini 3.5 Flash |
| Previous version | final_response_match_v2: 5 calls to Gemini 2.5 Flash |
Agent final response reference free
| Latest version | final_response_reference_free_v2 |
| Type | Adaptive rubrics |
| Description | An adaptive rubrics metric that evaluates the quality of an AI agent's final answer without needing a reference answer.
You need to provide rubrics for this metric, as it doesn't support auto-generated rubrics. |
| How to access in SDK | types.RubricMetric.FINAL_RESPONSE_REFERENCE_FREE |
| Input |
|
| Output |
|
| Number of LLM calls | 5 calls to Gemini 3.5 Flash |
| Previous version | final_response_reference_free_v1: 5 calls to Gemini 2.5 Flash |
Agent final response quality
| Latest version | final_response_quality_v2 |
| Type | Adaptive rubrics |
| Description | A comprehensive adaptive rubrics metric that evaluates the overall quality of an agent's response. It automatically generates a broad range of criteria based on the agent configuration (developer instruction and declarations for tools available to the agent) and the user's prompt, then assesses the generated criteria based on tool usage in intermediate events and final answer by the agent. |
| How to access in SDK | types.RubricMetric.FINAL_RESPONSE_QUALITY |
| Input |
|
| Output |
The score represents the passing rate of the response based on the rubrics. |
| Number of LLM calls | 5 calls to Gemini 3.5 Flash and 1 call to Gemini 3.1 Pro |
| Previous version | final_response_quality_v1: 5 calls to Gemini 2.5 Flash and 1 call to Gemini 2.5 Pro |
Agent hallucination
| Latest version | hallucination_v2 |
| Type | Static Rubrics |
| Description | A score-based metric that checks for factuality and consistency of text responses by segmenting the response into atomic claims. It verifies if each claim is grounded or not based on tool usage in the intermediate events.
It can also be leveraged to evaluate any intermediate text responses by setting the flag evaluate_intermediate_nl_responses to true.
|
| How to access in SDK | types.RubricMetric.HALLUCINATION |
| Input |
|
| Output |
0-1, and represents the ratio of sentences labeled as supported or no_rad to the total number of sentences.
The explanation field is a JSON string containing a list of per-event objects with the following schema:
[ { "response": "string", "score": "double", "explanation": [ { "sentence": "string", "label": "supported | unsupported | contradictory | disputed | no_rad", "rationale": "string", "supporting_excerpt": "string or null", "contradicting_excerpt": "string or null" } ] } ] explanation entry contains one object per segmented sentence with the following fields:
|
| Number of LLM calls | 2 calls to Gemini 3.5 Flash |
| Previous version | hallucination_v1: 2 calls to Gemini 2.5 Flash |
Agent tools usage quality
| Latest version | tool_use_quality_v2 |
| Type | Adaptive rubrics |
| Description | A targeted adaptive rubrics metric that evaluates the selection of appropriate tools, correct parameter usage, and adherence to the specified sequence of operations. |
| How to access in SDK | types.RubricMetric.TOOL_USE_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 5 calls to Gemini 3.5 Flash and 1 call to Gemini 3.1 Pro |
| Previous version | tool_use_quality_v1: 5 calls to Gemini 2.5 Flash and 1 call to Gemini 2.5 Pro |
Agent multi-turn task success
| Latest version | multi_turn_task_success_v1 |
| Type | Adaptive rubrics |
| Description |
An adaptive rubrics metric that evaluates whether the agent successfully fulfilled user goals across an entire multi-turn conversation. It focuses on observable outcomes and confirmations in the agent's responses rather than intermediate processes such as specific tool calls or reasoning steps.
The metric operates in three steps:
|
| How to access in SDK | types.RubricMetric.MULTI_TURN_TASK_SUCCESS |
| Input |
|
| Output |
|
| Number of LLM calls | 2 calls to Gemini 3.1 Pro and 5 calls to Gemini 3 Flash |
Agent multi-turn tool use quality
| Latest version | multi_turn_tool_use_quality_v1 |
| Type | Adaptive rubrics |
| Description |
An adaptive rubrics metric that evaluates the technical and semantic correctness of the agent's tool calls across an entire multi-turn conversation. It verifies that the agent selected the correct tools, populated arguments correctly, and adhered to the tool schemas for each user goal.
The metric operates in three steps:
|
| How to access in SDK | types.RubricMetric.MULTI_TURN_TOOL_USE_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 2 calls to Gemini 3.1 Pro and 5 calls to Gemini 3 Flash |
Agent multi-turn trajectory quality
| Latest version | multi_turn_trajectory_quality_v1 |
| Type | Adaptive rubrics |
| Description |
An adaptive rubrics metric that evaluates the quality of the agent's step-by-step execution trajectory across an entire multi-turn conversation. It focuses on the logical structure and technical validity of the agent's path rather than just the final response.
The metric operates in three steps:
|
| How to access in SDK | types.RubricMetric.MULTI_TURN_TRAJECTORY_QUALITY |
| Input |
|
| Output |
|
| Number of LLM calls | 2 calls to Gemini 3.1 Pro and 5 calls to Gemini 3 Flash |
Gecko text-to-image quality
| Latest version | gecko_text2image_v2 |
| Type | Adaptive rubrics |
| Description | The Gecko text-to-image metric is an adaptive, rubric-based method for evaluating the quality of a generated image against its corresponding text prompt. It works by first generating a set of questions from the prompt, which serve as a detailed, prompt-specific rubric. A model then answers these questions based on the generated image. |
| How to access in SDK | types.RubricMetric.GECKO_TEXT2IMAGE |
| Input |
|
| Output |
|
| Number of LLM calls | 2 calls to Gemini 3.5 Flash |
| Previous version | gecko_text2image_v1: 2 calls to Gemini 2.5 Flash |
Gecko text-to-video quality
| Latest version | gecko_text2video_v2 |
| Type | Adaptive rubrics |
| Description | The Gecko text-to-video metric is an adaptive, rubric-based method for evaluating the quality of a generated video against its corresponding text prompt. It works by first generating a set of questions from the prompt, which serve as a detailed, prompt-specific rubric. A model then answers these questions based on the generated video. |
| How to access in SDK | types.RubricMetric.GECKO_TEXT2VIDEO |
| Input |
|
| Output |
|
| Number of LLM calls | 2 calls to Gemini 3.5 Flash |
| Previous version | gecko_text2video_v1: 2 calls to Gemini 2.5 Flash |