Gemini 3.8 Flash-Lite TTS

Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) is Google's fast, cost-efficient text-to-speech model for high-throughput production workloads, available through Gemini Enterprise.

Capabilities

  • High-throughput efficiency: Built for bulk production, conversational voice agent cascades, read-aloud features, and everyday single-speaker speech in 101 languages.
  • Same schema as Gemini 3.8 Flash TTS: Uses the same request schema and prompting format as gemini-3.8-flash-tts, so you can switch models by changing the model ID.
  • Voice options: Works with 30 prebuilt voices, more than 2,000 curated voices in the Extended Voice Library, voices that you create with Voice design, and voices that you replicate with Voice replication.

For features, code samples, and prompting guidance, see Generate speech with Gemini TTS.

When to use which TTS model

Both Gemini 3.8 TTS models share the same request schema and prompting format. Choose the model that fits your workload:

Feature or workload Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) Gemini 3.8 Flash TTS (gemini-3.8-flash-tts)
Primary strength High throughput, low latency, and cost efficiency Voice fidelity, acting nuance, and dialect coverage
Best use cases High-volume production, real-time voice agent cascades, read-aloud features, voice replication, everyday single-speaker speech Audiobooks, studio narration, complex multi-speaker dialogue, frequent vocal-burst tags, difficult pronunciation, regional dialects
Supported languages 101 languages 130 languages

Pricing

Model ID gemini-3.8-flash-lite-tts
Modalities
Text Input only
Image Not supported
Audio Output only
Video Not supported
Token limits Input token limit 8,192
Output token limit 16,384
Capabilities
Consumption options
Supported regions

Model availability

  • Global: global
Versions
  • gemini-3.8-flash-lite-tts
    • Launch stage: Preview
    • Release date: September 28, 2026
Supported languages 101 languages. See Supported languages.

Get started

The following example generates single-speaker speech and saves it as a WAV file:

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-lite-tts",
    contents=[{
        "role": "user",
        "parts": [{
            "text": "Have a wonderful day!",
            "speech_metadata": {"style": "cheerful and friendly"},
        }],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
    },
)

# The SDK has already decoded the base64 audio, so inline_data.data is a
# complete WAV file by default.
with open("out.wav", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-lite-tts:generateContent \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [{
        "text": "Have a wonderful day!",
        "speechMetadata": {"style": "cheerful and friendly"}
      }]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.wav

If you use gemini-3.1-flash-tts-preview or another earlier Gemini TTS model, see the migration guide.