Generate speech with Gemini TTS

The Gemini text-to-speech (TTS) models convert text into single-speaker or multi-speaker audio. You can control the style, accent, pace, and tone of the audio with structured turn metadata (speech_metadata) and inline vocal tags.

This page shows you how to generate speech with the following models using the generateContent and streamGenerateContent methods of the Gemini Enterprise API:

The TTS models are tailored for exact text recitation with fine-grained control over style and sound, such as narration, audiobooks, and voice agent responses. For interactive, real-time conversations with audio input and output, use the Live API instead.

If you use an earlier Gemini TTS model, such as gemini-3.1-flash-tts-preview, see the migration guide.

Supported models

Model Single speaker Multi-speaker Streaming Voice design Voice replication
Gemini 3.8 Flash TTS (gemini-3.8-flash-tts)
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts)
Gemini 3.1 Flash TTS Preview (gemini-3.1-flash-tts-preview)
Gemini 2.5 Pro TTS (gemini-2.5-pro-tts)
Gemini 2.5 Flash TTS (gemini-2.5-flash-tts)

The samples and features on this page apply to the Gemini 3.8 TTS models. For the earlier models, see Gemini-TTS in the Google Cloud Text-to-Speech documentation.

When to use which model

Both Gemini 3.8 TTS models share the same request schema and prompting format, so you can switch between them by changing the model ID:

  • Use Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) when acoustic fidelity, nuanced acting, and expressive control are the top priority. It suits studio-grade creative work, complex multi-speaker dialogue, frequent vocal-burst tags, difficult pronunciations, regional or minority dialects, and long-form narration that needs a stable voice and room tone.
  • Use Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) for fast, cost-efficient production workloads. It's the recommended replacement for gemini-3.1-flash-tts-preview, and suits high-volume batch production, conversational voice agent cascades, read-aloud features, voice replication, and everyday single-speaker speech across major languages.

Before you begin

  1. Set up a project and enable the Gemini Enterprise API.
  2. Configure Application Default Credentials for your development environment.
  3. To use the Python samples, install version 2.25.0 or later of the Google Gen AI SDK:

    pip install --upgrade "google-genai>=2.25.0"
    

The Gemini 3.8 TTS models are available in the global location. Send requests to the aiplatform.googleapis.com endpoint.

Single-speaker TTS

To convert text to single-speaker audio, pass the verbatim transcript in parts[].text, add optional turn-level styling in parts[].speech_metadata, and set the voice in speechConfig.voiceConfig.voice. The voice field accepts a prebuilt voice name, an Extended Voice Library voice ID, or the ID (voice_...) of a designed or replicated voice (or an optional stateless voicekey_... key).

By default, unary requests return a complete WAV file (16-bit PCM, 24 kHz, mono), so you can write the audio bytes directly to a .wav file. Streaming requests return raw PCM chunks instead. For details and other encodings, see Audio output formats.

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-tts",
    contents=[{
        "role": "user",
        "parts": [{
            "text": "Have a wonderful day!",
            "speech_metadata": {"style": "cheerful and friendly"},
        }],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
    },
)

# The SDK has already decoded the base64 audio, so inline_data.data is a
# complete WAV file by default.
with open("out.wav", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [{
        "text": "Have a wonderful day!",
        "speechMetadata": {"style": "cheerful and friendly"}
      }]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.wav

Single-speaker streaming TTS

To receive audio while the model is still synthesizing it, use the streamGenerateContent method. Streaming responses return headerless raw 16-bit PCM chunks (24 kHz, mono), so you can pass each chunk to a player, a socket, or another audio pipeline as it arrives. In the following Python sample, the emit_audio function stands in for that destination.

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

def emit_audio(pcm: bytes) -> None:
    """Sends a chunk of raw 16-bit PCM audio (24 kHz, mono) downstream."""
    # Replace with your audio destination, such as a player, a WebSocket,
    # or a telephony stream.
    print(f"Received {len(pcm)} bytes of audio")

response_stream = client.models.generate_content_stream(
    model="gemini-3.8-flash-lite-tts",
    contents=[{
        "role": "user",
        "parts": [{
            "text": "Have a wonderful day!",
            "speech_metadata": {"style": "cheerful and friendly"},
        }],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
    },
)

for chunk in response_stream:
    if not chunk.candidates or not chunk.candidates[0].content:
        continue
    for part in chunk.candidates[0].content.parts or []:
        if part.inline_data and part.inline_data.data:
            emit_audio(part.inline_data.data)

REST

curl -N -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-lite-tts:streamGenerateContent?alt=sse" \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [{
        "text": "Have a wonderful day!",
        "speechMetadata": {"style": "cheerful and friendly"}
      }]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | sed -n 's/^data: //p' \
    | jq -r '.candidates[0].content.parts[0].inlineData.data // empty' \
    | while read -r chunk; do echo "$chunk" | base64 --decode; done > streamed.pcm

The streamed.pcm file contains raw 16-bit PCM audio at 24 kHz, mono.

Multi-speaker TTS

For dialogue between two speakers, configure both speakers in speechConfig.multiSpeakerVoiceConfig.speakerVoiceConfigs. Then pass each turn as a separate part, set speech_metadata.speaker on every turn to one of the configured speaker names, and add an optional turn-level style.

Multi-speaker requests require exactly two speakers.

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-tts",
    contents=[{
        "role": "user",
        "parts": [
            {
                "text": "How's it going today, Jane?",
                "speech_metadata": {"speaker": "Joe", "style": "cheerful and friendly"},
            },
            {
                "text": "Not too bad, how about you? Ready to test these new voices?",
                "speech_metadata": {"speaker": "Jane", "style": "calm and relaxed"},
            },
        ],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {
            "multi_speaker_voice_config": {
                "speaker_voice_configs": [
                    {"speaker": "Joe", "voice_config": {"voice": "Puck"}},
                    {"speaker": "Jane", "voice_config": {"voice": "Kore"}},
                ]
            }
        },
    },
)

with open("dialogue.wav", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [
        {
          "text": "How'\''s it going today, Jane?",
          "speechMetadata": {"speaker": "Joe", "style": "cheerful and friendly"}
        },
        {
          "text": "Not too bad, how about you? Ready to test these new voices?",
          "speechMetadata": {"speaker": "Jane", "style": "calm and relaxed"}
        }
      ]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "multiSpeakerVoiceConfig": {
          "speakerVoiceConfigs": [
            {"speaker": "Joe", "voiceConfig": {"voice": "Puck"}},
            {"speaker": "Jane", "voiceConfig": {"voice": "Kore"}}
          ]
        }
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > dialogue.wav

Multi-speaker streaming TTS

Multi-speaker requests also support streaming. Use the same request as in Multi-speaker TTS with the streamGenerateContent method.

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

def emit_audio(pcm: bytes) -> None:
    """Sends a chunk of raw 16-bit PCM audio (24 kHz, mono) downstream."""
    # Replace with your audio destination, such as a player, a WebSocket,
    # or a telephony stream.
    print(f"Received {len(pcm)} bytes of audio")

response_stream = client.models.generate_content_stream(
    model="gemini-3.8-flash-tts",
    contents=[{
        "role": "user",
        "parts": [
            {
                "text": "How's it going today, Jane?",
                "speech_metadata": {"speaker": "Joe", "style": "cheerful and friendly"},
            },
            {
                "text": "Not too bad, how about you? Ready to test these new voices?",
                "speech_metadata": {"speaker": "Jane", "style": "calm and relaxed"},
            },
        ],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {
            "multi_speaker_voice_config": {
                "speaker_voice_configs": [
                    {"speaker": "Joe", "voice_config": {"voice": "Puck"}},
                    {"speaker": "Jane", "voice_config": {"voice": "Kore"}},
                ]
            }
        },
    },
)

for chunk in response_stream:
    if not chunk.candidates or not chunk.candidates[0].content:
        continue
    for part in chunk.candidates[0].content.parts or []:
        if part.inline_data and part.inline_data.data:
            emit_audio(part.inline_data.data)

REST

curl -N -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:streamGenerateContent?alt=sse" \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [
        {
          "text": "How'\''s it going today, Jane?",
          "speechMetadata": {"speaker": "Joe", "style": "cheerful and friendly"}
        },
        {
          "text": "Not too bad, how about you? Ready to test these new voices?",
          "speechMetadata": {"speaker": "Jane", "style": "calm and relaxed"}
        }
      ]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "multiSpeakerVoiceConfig": {
          "speakerVoiceConfigs": [
            {"speaker": "Joe", "voiceConfig": {"voice": "Puck"}},
            {"speaker": "Jane", "voiceConfig": {"voice": "Kore"}}
          ]
        }
      }
    }
  }' | sed -n 's/^data: //p' \
    | jq -r '.candidates[0].content.parts[0].inlineData.data // empty' \
    | while read -r chunk; do echo "$chunk" | base64 --decode; done > dialogue_streamed.pcm

The dialogue_streamed.pcm file contains raw 16-bit PCM audio at 24 kHz, mono.

Control speech style with metadata and tags

The Gemini 3.8 TTS models treat the text field as a verbatim transcript. Instructions written into the text, such as "Say cheerfully: Hello!" or "Speaker 1: Hello!", might be read aloud. To control delivery, split your instructions by scope:

  • Sustained turn-level delivery (speech_metadata.style): Put emotion, delivery style, prosody, pacing, and volume that apply to a whole turn in speech_metadata.style. For example, "style": "whispered urgently", "style": "out of breath", or "style": "warm and enthusiastic".
  • Point-in-time events (inline tags): Put momentary non-speech vocal sounds and pauses inside the transcript in angle brackets. For example, "Wait... <short pause> did you hear that? <sigh>" or "Excuse me <cough> as I was saying...".

For more guidance, see the prompting guide.

Voice options

The Gemini 3.8 TTS models support four ways to select or create voices:

  1. Prebuilt voices: 30 curated voices.
  2. Extended Voice Library: More than 2,000 curated preset voices across languages, regional accents, and character personas.
  3. Voice design: Create a custom voice from a natural-language description with the Voices API create method (client.voices.create() in the Google Gen AI SDK, or POST https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices). Set type to VOICE_TYPE_PROMPTED and store to true. The response returns a stored voice_... ID and a sample of the voice.
  4. Voice replication: Replicate a speaker's voice from reference and consent audio with the same create method. Set type to VOICE_TYPE_REPLICATED. With store: true, the response returns a stored voice_... ID. With store: false, it returns a stateless voicekey_... key.

Custom voice limits and TTL

Voice type Storage mode Management Expiration
Stored voices (voice_..., designed or replicated) store: true Voices API list, get, and delete methods 1 year after last use
Voice replication keys (voicekey_...) store: false Managed by your application. Google doesn't retain the key. 7 days after creation

Generating speech with a stored voice restarts its one-year retention period, so a voice that you use regularly doesn't expire. Voice creation requests are subject to a per-project, per-minute quota. If you exceed it, the create method returns a RESOURCE_EXHAUSTED error. For more information, see Quotas and system limits.

Prebuilt voices

The following list shows each prebuilt voice name and its characteristic:

  • Achernar: Soft
  • Achird: Friendly
  • Algenib: Gravelly
  • Algieba: Smooth
  • Alnilam: Firm
  • Aoede: Breezy
  • Autonoe: Bright
  • Callirrhoe: Easy-going
  • Charon: Informative
  • Despina: Smooth
  • Enceladus: Breathy
  • Erinome: Clear
  • Fenrir: Excitable
  • Gacrux: Mature
  • Iapetus: Clear
  • Kore: Firm
  • Laomedeia: Upbeat
  • Leda: Youthful
  • Orus: Firm
  • Puck: Upbeat
  • Pulcherrima: Forward
  • Rasalgethi: Informative
  • Sadachbia: Lively
  • Sadaltager: Knowledgeable
  • Schedar: Even
  • Sulafat: Warm
  • Umbriel: Easy-going
  • Vindemiatrix: Gentle
  • Zephyr: Bright
  • Zubenelgenubi: Casual

Extended Voice Library

Beyond the 30 prebuilt voices, the Extended Voice Library provides more than 2,000 curated preset voices across languages, regional accents, character personas, and use cases such as audiobooks, conversational agents, and news. To use a library voice, pass its ID in speechConfig.voiceConfig.voice, the same way that you pass a prebuilt voice name.

To browse the library, call the Voices API list method (ListVoices). It returns the custom voices stored in your project, followed by the prebuilt and Extended Voice Library voices. Each library voice includes its id and metadata such as language_code, accent, gender, pitch, persona, and description. To narrow the results, use the following filters:

  • type: The voice source: prebuilt for prebuilt and library voices, prompted for designed voices, or replicated for replicated voices. If you pass more than one value, the method returns voices that match any of them.
  • search: Text to match, case-insensitively, against each voice's display name and description.

The following example searches the library for narrator voices:

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.voices.list(type_=["prebuilt"], search="narrator", page_size=50)
for voice in response.voices or []:
    print(voice.id, voice.language_code, voice.gender, voice.description)

REST

curl -G \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices \
  --data-urlencode "type=prebuilt" \
  --data-urlencode "search=narrator" \
  --data-urlencode "pageSize=50"

The list method returns up to 50 voices per page. To get the next page, pass the next_page_token value from the response as the page_token parameter.

Audio output formats

The default audio encoding depends on the method:

  • Unary requests (generateContent) return a complete WAV file (audio/wav) with a RIFF header: 16-bit PCM, 24 kHz, mono.
  • Streaming requests (streamGenerateContent) return headerless raw 16-bit signed little-endian PCM chunks (audio/l16) at 24 kHz, mono.

To request a different encoding, set generationConfig.responseFormat to an audio format. The Gemini 3.8 TTS models support the following mimeType values:

mimeType value Format Description
AUDIO_WAV (unary default) WAV (audio/wav) Complete WAV file with a RIFF header (16-bit PCM, 24 kHz, mono).
AUDIO_L16 (streaming default) Linear PCM (audio/l16) Headerless raw 16-bit signed little-endian PCM. Use for streaming, custom audio pipelines, or concatenating multi-turn clips.
AUDIO_MULAW μ-law (audio/mulaw) G.711 μ-law companded audio (8 kHz, mono), commonly used in North American and Japanese telephony.
AUDIO_ALAW A-law (audio/alaw) G.711 A-law companded audio (8 kHz, mono), commonly used in European and international telephony.

The models ignore the sampleRate field. Linear PCM and WAV output is 24 kHz, and μ-law and A-law output is 8 kHz, regardless of the rate in the response mimeType. Resample the output on the client if your pipeline requires a different sample rate.

The following example requests μ-law output for a telephony pipeline. The Google Gen AI SDK for Python doesn't expose responseFormat in GenerateContentConfig, so the Python sample passes it in the request body with http_options.extra_body.

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-tts",
    contents=[{"role": "user", "parts": [{"text": "Have a wonderful day!"}]}],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
        "http_options": {
            "extra_body": {
                "generationConfig": {
                    "responseFormat": [{"audio": {"mimeType": "AUDIO_MULAW"}}]
                }
            }
        },
    },
)

# The response is headerless μ-law audio at 8 kHz, mono.
with open("out.mulaw", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
  -d '{
    "contents": [{"role": "user", "parts": [{"text": "Have a wonderful day!"}]}],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "responseFormat": [{"audio": {"mimeType": "AUDIO_MULAW"}}],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.mulaw

Supported languages

The Gemini 3.8 TTS models detect the input language automatically. Gemini 3.8 Flash TTS supports 130 languages, and Gemini 3.8 Flash-Lite TTS supports 101 languages:

Limitations

  • The TTS models accept text-only input and return audio-only output.
  • Multi-speaker requests require exactly two speakers. To combine designed (voice_...) or replicated (voice_... or voicekey_...) voices in multi-character dialogue, or to use more than two speakers, synthesize each speaker's turn individually and concatenate the audio. Request AUDIO_L16 output for each turn, or strip the WAV header from each clip before you concatenate the clips.
  • The models are available only in the global location.
  • System instructions aren't supported.
  • You can't set the output sample rate. MP3 and Ogg Opus output aren't supported.

What's next