This page describes how to move from earlier Gemini TTS models, such as
gemini-3.1-flash-tts-preview, gemini-2.5-flash-tts, gemini-2.5-pro-tts,
and gemini-2.5-flash-lite-preview-tts, to
Gemini 3.8 Flash TTS
(gemini-3.8-flash-tts) and
Gemini 3.8 Flash-Lite TTS
(gemini-3.8-flash-lite-tts). The earlier models are documented in
Gemini-TTS in the Google Cloud
Text-to-Speech documentation.
Choose a replacement model
- Gemini 3.8 Flash-Lite TTS is the recommended replacement for
gemini-3.1-flash-tts-preview. Choose it for high-volume production, voice agents, read-aloud features, and everyday single-speaker speech. - Gemini 3.8 Flash TTS is the flagship model. Choose it when voice fidelity, acting nuance, multi-speaker dialogue, or dialect coverage matter most, such as for audiobooks and studio narration.
Both models use the same request schema, so you can switch between them by changing the model ID. For a comparison, see When to use which model.
Summary of changes
| Area | Earlier Gemini TTS models | Gemini 3.8 TTS models |
|---|---|---|
| API | Cloud Text-to-Speech API or Gemini Enterprise API | Gemini Enterprise API only (generateContent and streamGenerateContent) |
| Location | global and regional endpoints |
global only |
| Style direction | Written into the text ("Say the following in a curious way: ...") or the Cloud Text-to-Speech prompt field |
speech_metadata.style on each part |
| Multi-speaker dialogue | Speaker prefixes in one text block, or Cloud Text-to-Speech multiSpeakerMarkup |
One part per turn with speech_metadata.speaker |
| Inline vocal tags | Square brackets, such as [sigh] |
Angle brackets, such as <sigh> |
| Voice selection | prebuiltVoiceConfig.voiceName |
voiceConfig.voice, which also accepts designed voice IDs and voice replication keys. prebuiltVoiceConfig.voiceName still works for prebuilt voices. |
| Custom voices | Inline reference audio in replicatedVoiceConfig |
Voice design and voice replication through the Voices API |
| Default output (Gemini Enterprise API) | Raw 16-bit PCM, 24 kHz, mono | Unary requests: a complete WAV file (16-bit PCM, 24 kHz, mono). Streaming requests: raw 16-bit PCM, unchanged. You can request other encodings with responseFormat. |
temperature, topP, topK |
Ignored | Rejected with an INVALID_ARGUMENT error |
Update your requests
- Move style directions into
speech_metadata.style: The Gemini 3.8 TTS models readtextas a verbatim transcript, so directions written into the text might be spoken aloud. Put sustained acting, tone, prosody, and pacing directions inspeech_metadata.style. - Use one part per dialogue turn: For multi-speaker dialogue, pass each
turn as a separate
part, and setspeech_metadata.speakeron every turn to one of the speakers inmultiSpeakerVoiceConfig. - Replace square-bracket tags with angle-bracket tags: Use angle brackets,
such as
<sigh>,<laugh>, or<short pause>, only for point-in-time vocal events. Write disfluencies, such as "uhm", directly in the transcript. - Design personas upfront: Replace long "Audio Profile" or "Director's
Notes" prompts with a voice created in
Voice design,
and keep
styleshort or empty. - Remove unsupported parameters: Remove
temperature,topP,topK,candidateCount, andsystemInstructionfrom your requests. - Handle WAV output from unary requests: Unary requests now return a
complete WAV file instead of raw PCM. If your code adds a WAV header or
concatenates clips, remove the header step, or set
generationConfig.responseFormattoAUDIO_L16to keep raw PCM output. - Recreate replicated voices with the Voices API: Inline reference audio
in
replicatedVoiceConfigisn't supported. Create a stored voice or a voice replication key with the Voices API, and pass it invoiceConfig.voice. For details, see Voice replication.
For more guidance on writing prompts, see the prompting guide.
Example
The following example shows a gemini-3.1-flash-tts-preview request and the
equivalent Gemini 3.8 TTS request.
Before Gemini 3.8
from google import genai
from google.genai import types
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.1-flash-tts-preview",
contents="Say the following in a curious way: OK, so... tell me about this [uhm] AI thing.",
config=types.GenerateContentConfig(
response_modalities=["AUDIO"],
speech_config=types.SpeechConfig(
language_code="en-US",
voice_config=types.VoiceConfig(
prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name="Kore")
),
),
),
)
Gemini 3.8 or later
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.8-flash-lite-tts",
contents=[{
"role": "user",
"parts": [{
"text": "OK, so... uhm, tell me about this AI thing.",
"speech_metadata": {"style": "curious"},
}],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {
"language_code": "en-US",
"voice_config": {"voice": "Kore"},
},
},
)
Both responses return the audio in
response.candidates[0].content.parts[0].inline_data.data. The earlier model
returns raw 16-bit PCM audio (24 kHz, mono), and the Gemini 3.8 TTS model
returns a complete WAV file.
Migrate from the Cloud Text-to-Speech API
If you call Gemini TTS through the Cloud Text-to-Speech API
(texttospeech.googleapis.com), send your requests to the generateContent
method on aiplatform.googleapis.com instead. The following table maps the
Cloud Text-to-Speech request fields to the Gemini Enterprise API request fields:
Cloud Text-to-Speech API (text:synthesize) |
Gemini Enterprise API (generateContent) |
|---|---|
input.text |
contents[].parts[].text |
input.prompt |
contents[].parts[].speechMetadata.style |
input.multiSpeakerMarkup.turns[] (speaker, text) |
One part per turn, with text and speechMetadata.speaker |
voice.modelName |
The model ID in the request URL |
voice.name |
generationConfig.speechConfig.voiceConfig.voice |
voice.languageCode |
generationConfig.speechConfig.languageCode |
voice.multiSpeakerVoiceConfig.speakerVoiceConfigs[] (speakerAlias, speakerId) |
generationConfig.speechConfig.multiSpeakerVoiceConfig.speakerVoiceConfigs[] (speaker, voiceConfig.voice) |
audioConfig.audioEncoding: LINEAR16 |
generationConfig.responseFormat[].audio.mimeType: AUDIO_WAV (default for unary requests) |
audioConfig.audioEncoding: PCM |
AUDIO_L16 (default for streaming requests) |
audioConfig.audioEncoding: MULAW or ALAW |
AUDIO_MULAW or AUDIO_ALAW |
audioConfig.audioEncoding: MP3 or OGG_OPUS |
Not supported. Encode the audio on the client. |
audioConfig.sampleRateHertz |
Not supported. For output sample rates, see Audio output formats. |
The Cloud Text-to-Speech API returns base64-encoded audio in audioContent.
The Gemini Enterprise API returns it in
candidates[0].content.parts[0].inlineData.data.
What's next
- Try the samples in the Gemini TTS overview.
- Learn how to direct style and vocal sounds in the prompting guide.