Back to Features
Feature ๏ผ Gemini 3.8 Flash TTS

Gemi-chan gets a voice.

On September 22, 2026 (US time), Google made its speech generation model Gemini 3.8 Flash TTS generally available. You pick a voice, direct the performance, and can even create a brand-new voice just by describing it. The explainer video in this feature is voiced by Gemi-chan herself, which is to say, by Gemini 3.8 Flash TTS itself.

๐ŸŽ™
Video

First, hear it in Gemi-chan's voice

In about two minutes, Gemi-chan walks you through what it can do and how to use it (in Japanese). It has sound, so mind your volume.

Voice: Gemini 3.8 Flash TTS (voice Erinome) ๏ผ BGM: composed in Python code
Chapter 1

What launched? Two models

There are two: Flash TTS, built for expressiveness, and Flash-Lite TTS, built for speed and low cost. They succeed 3.1 Flash TTS, which came out in preview in April, and this time they are generally available (GA) from day one.

130 languages
Japanese included

Languages supported by Flash TTS. Flash-Lite TTS supports 101. Both support Japanese.

2,000+ voices
Voice count

Google's announced figure. The developer changelog says "150+," so the count depends on how you tally it.

~1.4 cents / min
One minute of speech

Calculated from Flash TTS output pricing (through December 31, 2026). There is also a free tier.

The expressive flagship

๐ŸŽญ Gemini 3.8 Flash TTS

  • Model ID gemini-3.8-flash-tts
  • Focused on acting and vocal nuance
  • 130 languages; strong on dialects and long readings
  • For narration, stories and video
The fast, affordable workhorse

โšก Gemini 3.8 Flash-Lite TTS

  • Model ID gemini-3.8-flash-lite-tts
  • Fast and cheap; replaces 3.1 Flash TTS
  • 101 languages
  • For voice agents, read-aloud and bulk jobs

Early reviews: a strong start.

Google says it ranked #1 (71.4) on Hume AI's voice design benchmark. Third-party Artificial Analysis was also reported to rank it #1 for pronunciation accuracy at 89.5% (as of September 23, 2026). Google also writes that it placed near the top in human listening tests across major languages, including Japanese. The first and last of these come from Google itself, though, so your own ears are the best judge.

Chapter 2

Choosing a voice: four levels

From "pick a ready-made voice" to "make your own," there are four levels. The easiest is โ‘ ; the newest this time is โ‘ข.

30 preset voices

Just name one

Thirty voices with personality notes, such as Kore (firm), Puck (upbeat), Leda (youthful) and Erinome (clear). Start here.

Voice library

Search by language, gender, pitch

A much bigger shelf of voices. Filter by language, gender, pitch and persona to find one.

Voice design

Describe it in words

Describe something like "a crisp, energetic sports announcer in her 30s," and it creates and saves a new voice. You call it again by its ID, so a character's voice stays consistent.

Voice replication

Requires the speaker's consent

Recreates a person's voice from 10 to 30 seconds of audio. A recording of the same person reading a consent statement is required, so you cannot simply use someone else's voice.

Chapter 3

Direct the performance with "style" and tags

The same voice can sound cheerful, whispery or serious depending on your direction. There are three places to put that direction.

๐ŸŽญ

Delivery goes in "style"

"Whispering," "a little serious," "explaining slowly": the delivery for the whole line. It goes in a field separate from the transcript. It also works when written in Japanese (the video in this feature was directed in Japanese).

๐Ÿ˜†

Momentary sounds use tags

Put <laugh>, <sigh>, <short pause> and the like inside the transcript. Keep the tags in English even in a Japanese script.

๐Ÿ‘ญ

Two-speaker dialogue

Set up to two speakers and generate a conversation in one pass. In 3.8, every line must say who is speaking (preset voices only).

๐Ÿ” 

Emphasis with capitals

In an English script, writing a word in capitals gives it emphasis.

Gemi-chan (Gemini)
Casting

Gemi-chan's voice in the video

gemini-3.8-flash-tts Voice Erinome (clear) 17 lines

The video above pairs the preset voice Erinome with the delivery directions below. For scenes like whispering, showing off or saying goodbye, we added a little more direction line by line.

๐ŸŽ€ Everyday delivery

"Super cute, with a sweet, sparkly voice, bright and energetic, bouncy, a little fast" (written in Japanese)

๐Ÿคซ The whisper scene

"In a whisper, like telling a secret, super cute." The voice gets very quiet, so we raised its volume a little in the video.

Chapter 4

What changed from 3.1

Scripts and programs written for the previous model may not behave as expected as they are.

Up to 3.1 Flash TTS

๐Ÿ“ Directions could sit in the script

  • You could start a line with "Say cheerfully:"
  • Long "director's notes" set the voice's personality
  • Output was raw audio with no header
  • Called with generateContent
From 3.8 Flash TTS

๐ŸŽฏ The script is read exactly as written

  • Directions in the text get read aloud; put them in style
  • Create the voice's personality first with voice design
  • Returns a WAV file (24 kHz, mono)
  • Called through the new Interactions API
Chapter 5

Try it: in the browser or in code

To just try it, the browser is enough. To build it into your own app or video workflow, use the API.

STEP 1 ๐ŸŒ

Open Google AI Studio

Go to the speech generation page (aistudio.google.com/generate-speech). With a Google account, you can try it on the free tier.

STEP 2 ๐ŸŽค

Pick or make a voice

Choose from the presets or the voice library, or use voice design to create one from a description.

STEP 3 โ–ถ๏ธ

Add a script and play

Enter the script and delivery, then play. Copy a voice's voice_โ€ฆ ID to call the same voice from the API.

API

The API is a "text, style, voice" set

Create a Gemini API key in AI Studio, then send the text to read, the delivery (style) and a voice name (or voice_โ€ฆ ID). The audio comes back as base64 text, so decode it and save it as a .wav file.

Try it once from the terminal (curl)
โ‘  Set your key (the one you created in AI Studio)
$ export GEMINI_API_KEY="your-key"
โ‘ก Send text, style and voice, and save to out.wav (requires jq)
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \ -H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \ -d '{"model":"gemini-3.8-flash-tts","input":[{"type":"user_input","content":[{"type":"text","text":"Hi! I am Gemi-chan.","annotations":[{"type":"speech_metadata","style":"bright and cheerful"}]}]}],"response_format":{"type":"audio"},"generation_config":{"speech_config":[{"voice":"Erinome"}]}}' \ | jq -r '[.steps[] | select(.type=="model_output") | .content[] | select(.type=="audio")] | last | .data' \ | base64 --decode > out.wav
โœ“ Gemi-chan's voice is saved to out.wav

โ€ป Based on the official Gemini API documentation (ai.google.dev) as of September 26, 2026.
Keep your key private. Rather than writing it into a page or program, pass it in through an environment variable or similar.

Using Python (google-genai 2.25.0 or later)
Install the library
$ pip install -U google-genai
Save as tts.py and run it
import base64 from google import genai client = genai.Client() res = client.interactions.create( model="gemini-3.8-flash-tts", input=[{"type": "user_input", "content": [{ "type": "text", "text": "Hi! I am Gemi-chan.", "annotations": [{"type": "speech_metadata", "style": "bright and cheerful"}], }]}], response_format={"type": "audio"}, generation_config={"speech_config": [{"voice": "Erinome"}]}, ) with open("out.wav", "wb") as f: f.write(base64.b64decode(res.output_audio.data))

โ€ป This is different from the generate_content-plus-hand-made-WAV-header approach in guides written for 3.1 and earlier. Old code won't work as is.

Chapter 6

Pricing: discounted through year-end

Prices are per million tokens. Audio is 25 tokens per second, so one minute of audio is 1,500 tokens. Current prices run through December 31, 2026 and double on January 1, 2027.

Model Input (text) Output (audio) Note
3.8 Flash TTS
The expressive flagship
$0.50 $9 About 1.4 cents a minute. From 2027: $1 ๏ผ $18.
3.8 Flash-Lite TTS
Speed and savings
$0.50 $6 About 0.9 cents a minute. From 2027: $1 ๏ผ $12.
Batch
Jobs that can wait
Half Half For Flash TTS, $0.25 ๏ผ $4.50.
3.1 Flash TTS
The earlier preview
$1 $20 The new 3.8 is far cheaper through year-end.

โ€ป Prices are from the official Gemini API pricing page as of September 26, 2026, all per million tokens.
There is a free tier, but what you send on the free tier is used to improve Google's products (not on the paid tier).

Chapter 7

Good to know before you start

โš ๏ธ Five things to keep in mind

  • Voice replication requires consent. You need a recording of the same adult speaker reading a fixed consent statement. If the two recordings use different microphones, the speaker check can fail.
  • Generated audio carries a watermark. Audio from Google's Gemini Audio models includes SynthID, a mark showing it was made by AI. You can't hear it, but it can be detected later.
  • Only two speakers per dialogue. And only with preset voices. To make custom voices talk to each other, generate each line separately and join them.
  • Not yet on Google Cloud for businesses (Vertex AI). As of September 26, 2026, it is available through the Gemini API and AI Studio. Gemini Enterprise is listed as "coming soon."
  • The way you write scripts has changed. A script for the old model that starts with "Say cheerfully:" will have that direction read aloud. Move directions into the style field.
Choose โ†’ Direct โ†’ Design

The head of the broadcasting club finally got a real microphone.

Choose a voice, direct the performance, and if nothing fits, describe a new voice in words. Work that used to need a recording studio and a voice actor now starts somewhere anyone who can write a script can reach. If you can write the script, the voice can follow. The head of the broadcasting club has just started handing out the tools to get in the door. Start by finding one voice you like in AI Studio.

Related

Read next