Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) is Google's flagship
text-to-speech model, available through Gemini Enterprise. It's built for
studio-grade voice fidelity, expressive acting, authentic regional accents, and
stable long-form, multi-turn audio.
Capabilities
- Acoustic fidelity and acting nuance: Delivers a wide emotional range,
natural cadence, and close adherence to turn-level
styledirections and inline vocal events such as<laugh>,<sigh>, and<short pause>. - Long-form, multi-turn stability: Keeps voice identity, timbre, volume, and room tone consistent across long dialogues and multi-minute narration.
- Regional accents and pronunciation: Supports regional accents and minority dialects in 130 languages.
- Voice options: Works with 30 prebuilt voices, more than 2,000 curated voices in the Extended Voice Library, voices that you create with Voice design, and voices that you replicate with Voice replication.
For features, code samples, and prompting guidance, see Generate speech with Gemini TTS.
When to use which TTS model
Both Gemini 3.8 TTS models share the same request schema and prompting format. Choose the model that fits your workload:
| Feature or workload | Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) |
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) |
|---|---|---|
| Primary strength | Voice fidelity, acting nuance, and dialect coverage | High throughput, low latency, and cost efficiency |
| Best use cases | Audiobooks, studio narration, complex multi-speaker dialogue, frequent vocal-burst tags, difficult pronunciation, regional dialects | High-volume production, real-time voice agent cascades, read-aloud features, voice replication, everyday single-speaker speech |
| Supported languages | 130 languages | 101 languages |
| Model ID | gemini-3.8-flash-tts |
|
|---|---|---|
| Modalities |
|
|
| Token limits | Input token limit | 8,192 |
| Output token limit | 16,384 | |
| Capabilities |
|
|
| Consumption options |
|
|
| Supported regions |
|
|
| Versions |
|
|
| Supported languages | 130 languages. See Supported languages. | |
Get started
The following example generates single-speaker speech and saves it as a WAV file:
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.8-flash-tts",
contents=[{
"role": "user",
"parts": [{
"text": "Have a wonderful day!",
"speech_metadata": {"style": "cheerful and friendly"},
}],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {"voice_config": {"voice": "Kore"}},
},
)
# The SDK has already decoded the base64 audio, so inline_data.data is a
# complete WAV file by default.
with open("out.wav", "wb") as f:
f.write(response.candidates[0].content.parts[0].inline_data.data)
REST
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
-d '{
"contents": [{
"role": "user",
"parts": [{
"text": "Have a wonderful day!",
"speechMetadata": {"style": "cheerful and friendly"}
}]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"voiceConfig": {"voice": "Kore"}
}
}
}' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.wav
If you use an earlier Gemini TTS model, see the migration guide.