On September 22, 2026 (US time), Google made its speech generation model Gemini 3.8 Flash TTS generally available. You pick a voice, direct the performance, and can even create a brand-new voice just by describing it. The explainer video in this feature is voiced by Gemi-chan herself, which is to say, by Gemini 3.8 Flash TTS itself.
In about two minutes, Gemi-chan walks you through what it can do and how to use it (in Japanese). It has sound, so mind your volume.
There are two: Flash TTS, built for expressiveness, and Flash-Lite TTS, built for speed and low cost. They succeed 3.1 Flash TTS, which came out in preview in April, and this time they are generally available (GA) from day one.
Languages supported by Flash TTS. Flash-Lite TTS supports 101. Both support Japanese.
Google's announced figure. The developer changelog says "150+," so the count depends on how you tally it.
Calculated from Flash TTS output pricing (through December 31, 2026). There is also a free tier.
Google says it ranked #1 (71.4) on Hume AI's voice design benchmark. Third-party Artificial Analysis was also reported to rank it #1 for pronunciation accuracy at 89.5% (as of September 23, 2026). Google also writes that it placed near the top in human listening tests across major languages, including Japanese. The first and last of these come from Google itself, though, so your own ears are the best judge.
From "pick a ready-made voice" to "make your own," there are four levels. The easiest is โ ; the newest this time is โข.
Thirty voices with personality notes, such as Kore (firm), Puck (upbeat), Leda (youthful) and Erinome (clear). Start here.
A much bigger shelf of voices. Filter by language, gender, pitch and persona to find one.
Describe something like "a crisp, energetic sports announcer in her 30s," and it creates and saves a new voice. You call it again by its ID, so a character's voice stays consistent.
Recreates a person's voice from 10 to 30 seconds of audio. A recording of the same person reading a consent statement is required, so you cannot simply use someone else's voice.
The same voice can sound cheerful, whispery or serious depending on your direction. There are three places to put that direction.
"Whispering," "a little serious," "explaining slowly": the delivery for the whole line. It goes in a field separate from the transcript. It also works when written in Japanese (the video in this feature was directed in Japanese).
Put <laugh>, <sigh>, <short pause> and the like inside the transcript. Keep the tags in English even in a Japanese script.
Set up to two speakers and generate a conversation in one pass. In 3.8, every line must say who is speaking (preset voices only).
In an English script, writing a word in capitals gives it emphasis.
The video above pairs the preset voice Erinome with the delivery directions below. For scenes like whispering, showing off or saying goodbye, we added a little more direction line by line.
"Super cute, with a sweet, sparkly voice, bright and energetic, bouncy, a little fast" (written in Japanese)
"In a whisper, like telling a secret, super cute." The voice gets very quiet, so we raised its volume a little in the video.
Scripts and programs written for the previous model may not behave as expected as they are.
To just try it, the browser is enough. To build it into your own app or video workflow, use the API.
Go to the speech generation page (aistudio.google.com/generate-speech). With a Google account, you can try it on the free tier.
Choose from the presets or the voice library, or use voice design to create one from a description.
Enter the script and delivery, then play. Copy a voice's voice_โฆ ID to call the same voice from the API.
Create a Gemini API key in AI Studio, then send the text to read, the delivery (style) and a voice name (or voice_โฆ ID). The audio comes back as base64 text, so decode it and save it as a .wav file.
$ export GEMINI_API_KEY="your-key"curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"gemini-3.8-flash-tts","input":[{"type":"user_input","content":[{"type":"text","text":"Hi! I am Gemi-chan.","annotations":[{"type":"speech_metadata","style":"bright and cheerful"}]}]}],"response_format":{"type":"audio"},"generation_config":{"speech_config":[{"voice":"Erinome"}]}}' \
| jq -r '[.steps[] | select(.type=="model_output") | .content[] | select(.type=="audio")] | last | .data' \
| base64 --decode > out.wav
โป Based on the official Gemini API documentation (ai.google.dev) as of September 26, 2026.
Keep your key private. Rather than writing it into a page or program, pass it in through an environment variable or similar.
$ pip install -U google-genaiimport base64
from google import genai
client = genai.Client()
res = client.interactions.create(
model="gemini-3.8-flash-tts",
input=[{"type": "user_input", "content": [{
"type": "text", "text": "Hi! I am Gemi-chan.",
"annotations": [{"type": "speech_metadata", "style": "bright and cheerful"}],
}]}],
response_format={"type": "audio"},
generation_config={"speech_config": [{"voice": "Erinome"}]},
)
with open("out.wav", "wb") as f:
f.write(base64.b64decode(res.output_audio.data))โป This is different from the generate_content-plus-hand-made-WAV-header approach in guides written for 3.1 and earlier. Old code won't work as is.
Prices are per million tokens. Audio is 25 tokens per second, so one minute of audio is 1,500 tokens. Current prices run through December 31, 2026 and double on January 1, 2027.
| Model | Input (text) | Output (audio) | Note |
|---|---|---|---|
| 3.8 Flash TTS The expressive flagship |
$0.50 | $9 | About 1.4 cents a minute. From 2027: $1 ๏ผ $18. |
| 3.8 Flash-Lite TTS Speed and savings |
$0.50 | $6 | About 0.9 cents a minute. From 2027: $1 ๏ผ $12. |
| Batch Jobs that can wait |
Half | Half | For Flash TTS, $0.25 ๏ผ $4.50. |
| 3.1 Flash TTS The earlier preview |
$1 | $20 | The new 3.8 is far cheaper through year-end. |
โป Prices are from the official Gemini API pricing page as of September 26, 2026, all per million tokens.
There is a free tier, but what you send on the free tier is used to improve Google's products (not on the paid tier).
Choose a voice, direct the performance, and if nothing fits, describe a new voice in words. Work that used to need a recording studio and a voice actor now starts somewhere anyone who can write a script can reach. If you can write the script, the voice can follow. The head of the broadcasting club has just started handing out the tools to get in the door. Start by finding one voice you like in AI Studio.