This page describes best practices for writing prompts when using Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
For request syntax and code samples, see the Gemini TTS overview.
The Gemini 3.8 TTS models treat the text field as a verbatim transcript.
Sustained, turn-level direction goes in speech_metadata, and point-in-time
vocal events go inline in the transcript. The following request combines a
turn-level style with inline vocal tags and pauses:
{
"contents": [{
"role": "user",
"parts": [{
"text": "Wait... <short pause> did you hear that? <gasp> Someone's at the door.",
"speechMetadata": {"style": "whispered, nervous"}
}]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {"voiceConfig": {"voice": "Kore"}}
}
}
Style field versus inline tags
Split your performance instructions by scope:
- Turn-level delivery (
speech_metadata.style): Put sustained delivery attributes, such as emotion, prosody, overall pace, or delivery style (like"whispering","out of breath","muttering", or"sarcastic"), into thestylefield ofspeech_metadata. To create a stable character and performance across turns, design the persona upfront in Voice design and usestyleonly for optional turn-level tweaks. - Point-in-time events (inline tags): Put momentary non-speech vocal
bursts, breaths, or pauses inline inside the transcript using angle brackets
(
<cough>,<breath>,<sigh>,<short pause>). Use angle brackets (<...>) for highest audio quality, and stick to human vocalizations rather than non-vocal sound effects.
| Scope | Where to place | Examples |
|---|---|---|
| Turn-level (sustained across the turn) | speech_metadata.style |
"angry tone", "speaking rapidly", "out of breath", "whispers", "sarcastic" |
| Point-in-time (occurs at a specific word) | Inline in text (<...>) |
"<cough> Thank you all for coming tonight! <throat-clearing> As I was saying..." |
Pacing and pauses
You can control rhythm and silence at three levels:
- Punctuation and ellipses: Use commas, dashes (
--), and ellipses (...) for natural hesitation. - Inline pause tags: Insert
<short pause>or<long pause>where the speaker should pause. For example:"Hold on, let me think... <short pause> Alright, I've got it." - Turn-level pace: Set
"style": "speaking rapidly"or"style": "speaking slowly"inspeech_metadatato control the speaking rate for the whole turn.
Prosody and pitch
Use speech_metadata.style to control prosody, pitch, and inflection across a
turn. For example, "style": "high pitch, cheerful and excited inflection" or
"style": "monotone and flat". If the emotion shifts mid-dialogue, split the
script into separate turns, each with its own style.
Emphasis
Capitalize words in the transcript, together with punctuation and inline tags,
to stress them. For example:
"This is a VERY important point!" or
"It was a VERY long day <sigh> ... nobody listens anymore."
Vocal bursts and non-speech sounds
Place non-speech human vocalizations inline in angle brackets (<...>) at the
point where the sound should occur. Use human vocalizations rather than
non-vocal sound effects such as applause. Recommended vocal tags include:
| Tag | Tag | Tag | Tag |
|---|---|---|---|
<argh> |
<breath> |
<heavy breath> |
<exhales> |
<cackle> |
<cheer> |
<chuckle> / <chuckles> |
<cough> |
<cry> |
<gasp> |
<giggle> |
<groan> |
<growl> |
<grunt> |
<grr> |
<hiss> |
<laugh> / <laughter> |
<moan> |
<pant> |
<pff> / <phew> |
<scream> |
<shout> |
<shriek> |
<sigh> / <sighs> |
<sneeze> |
<snicker> |
<snort> |
<sob> |
<throat-clearing> |
<tsk> |
<whimper> |
<whispers> / <whispering> |
<yawn> |
<short pause> |
<long pause> |
Backchannels and overlapping speech
In multi-speaker dialogue, wrap listener reactions in pipe characters
(|reaction|) inside the active speaker's turn. This creates natural
backchannels or overlapping speech without a separate turn for each reaction.
- Short backchannels: Layer brief listener reactions inside the active
speaker's turn:
- Turn 1 (Speaker A):
"So the launch is Thursday |oh hmm| Are we actually ready?" - Turn 2 (Speaker B):
"Ready enough |oh really?| The last blocker cleared this morning." - Turn 3 (Speaker A):
"Then let's ship it |absolutely| and watch the dashboards."
- Turn 1 (Speaker A):
- Overlapping and interleaved speech: Use multiple pipe segments to
simulate simultaneous speech. This works best with
gemini-3.8-flash-tts:- Countdown or chorus:
"Let's surprise him on three |ok| ready?"followed by"one. two. three. |happy| happy |birthday| birthday!" - Full overlap:
"Hello |oh| there |my| it |goodness| must |gracious| be |would| almost |you| time |look| for |at that| dinner"
- Countdown or chorus:
Consistency across generations and what to avoid
Follow these guidelines to keep vocal identity stable across turns:
- Design personas upfront in Voice design instead of long style blocks:
Long-form
"Audio Profile"paragraphs and multi-bullet"Director's Notes"carried over from earlier models are the most common cause of voice drift. Use that same creative intuition upfront in Voice design to generate a persistent customvoice_...persona, then carry that voice ID through your TTS calls. - Rely on the voice reference for stability (omit meta-instructions):
Gemini 3.8 TTS models are trained to anchor on the audio reference first.
Don't include instructions telling the model to hold the voice steady (such
as
"do not switch speaker identity"or"maintain identical timbre"). Extra prompt text increases drift. Drop unnecessary style instructions and let the model vary naturally around the stable point provided by the voice reference. - Don't try to change immutable speaker traits in
style: Avoid putting age, gender, names, or permanent accent changes inspeech_metadata.style. Instead, pick a regional voice from the Extended Voice Library or create one with Voice design.
Recommended workflow
- Build the character once: Create your character in Voice design or select a regional voice from the Extended Voice Library that matches your target language and persona.
- Write natural spoken transcripts with disfluencies: For maximum
naturalness, write the
textas a real spoken transcript, including natural conversational disfluencies and hesitations (for example,"Oh uh yeah I think... hm, so that's interesting"). - Test plain TTS first: Synthesize your transcript with an empty
stylefield first. Most requests need nostyleinstruction at all. - Add short
styleprompts only for tweaks: Add a concisestylestring (such as"casual, friendly"or"muttering, then reassuring") only for turns that need a specific delivery adjustment, and reuse that exact short string across turns when you want a consistent baseline.
Multi-turn dialogue and voice agents
When you build a conversational voice agent or another multi-turn application:
- Make one TTS request for each turn as the LLM text arrives.
- Let the configured
voicecarry the speaker's identity across turns. Don't resend a long persona description on each turn. - Leave
styleempty, or send one short constant string for the whole conversation. - Split long agent responses into shorter turns rather than strengthening the
styleprompt.
What's next
- Create a consistent character with Voice design.
- Update prompts written for earlier models with the migration guide.