You can create a custom voice from a natural-language description using the
voice design feature. Instead of choosing a prebuilt voice or recording
reference audio, you create a voice by describing a character's age, vocal
timbre, accent, and baseline delivery. The Voices API stores the voice in your
project and returns a reusable voice_... ID that you pass in speech
generation requests.
Designed voices work with both
Gemini 3.8 Flash TTS
(gemini-3.8-flash-tts) and
Gemini 3.8 Flash-Lite TTS
(gemini-3.8-flash-lite-tts).
How voice design works
- Create a prompted voice: Call the Voices API
createmethod withvoice.typeset toVOICE_TYPE_PROMPTEDandstoreset totrue. - Receive a voice ID and a sample: The Voices API generates the voice,
stores it in your project, and returns its ID (for example,
voice_abc123...) in theidfield and a sample of the voice in thesample_audiofield. - Generate speech: Pass the voice ID in
speechConfig.voiceConfig.voicewhen you callgenerateContentorstreamGenerateContent.
Before you begin
Complete the steps in
Before you begin
on the Gemini TTS page. The Voices API is available in the global location
through the v1beta1 API version.
Create a designed voice
Set the following fields in the request:
store: Must betrue. Designed voices are always stored in your project.voice.type:VOICE_TYPE_PROMPTED.voice.prompted.input: A natural-language description of the voice.voice.display_name,voice.gender,voice.language_code: Optional metadata for the voice.
Don't set voice.model. Requests that set a model together with
store: true fail with an INVALID_ARGUMENT error.
Designing a voice can take several seconds. The Python sample sets the request timeout to 60 seconds.
The response contains the following fields:
id: The ID of the new voice.sample_audio: A sample of the voice. Thedatafield holds the base64-encoded audio, and themime_typefield names its encoding. For example,audio/l16; rate=24000; channels=1is headerless 16-bit PCM at 24 kHz, mono.usage: The number of tokens that the request used. For more information, see Pricing.
Python
import base64
import wave
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
voice = client.voices.create(
store=True,
voice={
"type": "VOICE_TYPE_PROMPTED",
"display_name": "Warm British Astronomer",
"gender": "male",
"language_code": "en-GB",
"prompted": {
"input": (
"A warm, thoughtful astronomer in his late 60s with a gentle"
" British accent, speaking with quiet wonder."
)
},
},
timeout=60,
)
print(f"Created voice ID: {voice.id}")
# Save the voice sample. The SDK returns sample_audio.data as a base64
# string. Raw 16-bit PCM (audio/l16) needs a WAV header.
sample = base64.b64decode(voice.sample_audio.data)
if voice.sample_audio.mime_type.lower().startswith("audio/l16"):
with wave.open("voice_sample.wav", "wb") as wf:
wf.setnchannels(1)
wf.setsampwidth(2)
wf.setframerate(24000)
wf.writeframes(sample)
else:
with open("voice_sample.wav", "wb") as f:
f.write(sample)
REST
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices \
-d '{
"store": true,
"voice": {
"type": "VOICE_TYPE_PROMPTED",
"displayName": "Warm British Astronomer",
"gender": "male",
"languageCode": "en-GB",
"prompted": {
"input": "A warm, thoughtful astronomer in his late 60s with a gentle British accent, speaking with quiet wonder."
}
}
}' > created_voice.json
# Print the voice ID and the encoding of the voice sample.
jq -r '.id, .sample_audio.mime_type' created_voice.json
# Decode the voice sample.
jq -r '.sample_audio.data' created_voice.json | base64 --decode > voice_sample.pcm
If sample_audio.mime_type is audio/l16, the voice_sample.pcm file
contains headerless 16-bit PCM audio at 24 kHz, mono.
To hear the voice with your own text, generate speech with it as shown in the next section.
Generate speech with a designed voice
Pass the voice ID in speechConfig.voiceConfig.voice:
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.8-flash-tts",
contents=[{
"role": "user",
"parts": [{
"text": (
"Look out past the rings of Saturn. Those faint photons left"
" their source millions of years ago."
),
"speech_metadata": {"style": "reflective and awe-inspired"},
}],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {"voice_config": {"voice": "VOICE_ID"}},
},
)
# The SDK has already decoded the base64 audio, so inline_data.data is a
# complete WAV file by default.
with open("designed_voice.wav", "wb") as f:
f.write(response.candidates[0].content.parts[0].inline_data.data)
REST
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
-d '{
"contents": [{
"role": "user",
"parts": [{
"text": "Look out past the rings of Saturn. Those faint photons left their source millions of years ago.",
"speechMetadata": {"style": "reflective and awe-inspired"}
}]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"voiceConfig": {"voice": "VOICE_ID"}
}
}
}' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > designed_voice.wav
Replace VOICE_ID with the voice ID that the create method
returned.
Manage stored voices
You can list, get, and delete the voices stored in your project, including
designed voices and
stored replicated voices.
The list method returns the voices stored in your project, followed by the
prebuilt and
Extended Voice Library
voices. To list only the voices stored in your project, set the type filter
to prompted and replicated, as the following samples do. The get method
also returns the sample of a designed voice in the sample_audio field.
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
# List the voices stored in your project.
for voice in client.voices.list(type_=["prompted", "replicated"]).voices or []:
print(voice.id, voice.display_name, voice.type)
# Get a voice by ID.
voice = client.voices.get(id="VOICE_ID")
# Delete a voice.
client.voices.delete(id="VOICE_ID")
REST
# List the voices stored in your project.
curl -G \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices \
--data-urlencode "type=prompted" \
--data-urlencode "type=replicated"
# Get a voice by ID.
curl -H "Authorization: Bearer $(gcloud auth print-access-token)" \
https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices/VOICE_ID
# Delete a voice.
curl -X DELETE \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices/VOICE_ID
Limits
- Stored voices expire one year after they were last used. Generating speech with a voice restarts this period, so a voice that you use regularly doesn't expire. Delete voices that you no longer need.
- Voice creation requests are subject to a per-project, per-minute quota. If
you exceed it, the
createmethod returns aRESOURCE_EXHAUSTEDerror. Wait and retry. For more information, see Quotas and system limits.
Pricing
Creating a designed voice is billed as text input tokens and audio output
tokens at the Gemini 3.8 Flash TTS rates. The audio output tokens cover the
voice sample that the Voices API generates. The usage field in the create
response reports both counts. For example, the one-sentence description in the
create sample uses about 220 input tokens and about 1,100 output
tokens.
The get, list, and delete methods aren't billed. Speech that you generate
with a designed voice is billed at the rates of the model that you call.
For token rates, see Gemini Enterprise pricing.
Prompting best practices
- Put permanent traits in the voice description: Define age, gender,
timbre, vocal texture, and regional accent when you create the voice, not
in
speech_metadata.style. - Use
speech_metadata.stylefor situational emotion: After you create the voice, use shortstylevalues, such as"whispered urgently"or"cheerful and energetic", to direct each turn without changing the speaker's identity. - Be specific and concise: A clear one- or two-sentence description, such as "A crisp, energetic sports announcer in her 30s with a slight Midwestern accent", produces more consistent results than a long or contradictory paragraph.
What's next
- Replicate an existing speaker's voice with Voice replication.
- Learn about turn-level styling, inline tags, and multi-speaker dialogue in Generate speech with Gemini TTS.