Gemini 3.8 Speech Models: A UK SME Guide
Gemini 3.8 speech models are Google DeepMind's newest text-to-speech systems, Gemini 3.8 Flash TTS and Flash-Lite TTS, launched on 23 September 2026. They add script-level control over tone, pacing, and pronunciation across 100-plus languages. Voice cloning remains restricted in the UK, but UK SMEs can still design custom voices and direct delivery line by line.
Gemini 3.8 Speech Models: Key Facts at a Glance
- Two models, two jobs — Flash TTS is built for creative direction (character voices, audiobooks, podcasts); Flash-Lite TTS is built for high-volume, cost-efficient narration (dubbing, voice agents, bulk content) (Google, 2026).
- Delivery is scriptable — both models follow line-by-line stage directions for tone, pacing, and emphasis, plus vocal bursts such as
<laughs>and<sigh>(Google, 2026). - Voice cloning is off in the UK — replication of a real person's voice through Google AI Studio is unavailable in the UK, the EEA, Switzerland, India, Texas, and Illinois (Android Authority, 2026).
- Pricing is per token, not per minute — Flash-Lite TTS runs at $0.50 per million input tokens and $6.00 per million audio output tokens until 31 December 2026 (Google AI for Developers, 2026).
- Every clip is watermarked — SynthID and C2PA credentials are embedded in all generated audio, regardless of region (Google, 2026).
What Are Gemini 3.8 Flash TTS and Flash-Lite TTS?
Google DeepMind announced Gemini 3.8 Flash TTS and Flash-Lite TTS on 23 September 2026, rolling both out the same day through the Gemini API and Google AI Studio (Google, 2026).
The two models split by job. Flash TTS is aimed at "deep creative direction and character design across gaming, immersive audiobooks, podcasts, and interactive media." Flash-Lite TTS is aimed at "high-volume, cost-efficient use, optimized for dubbing, audio content creation, and voice agents" (Unite.AI, 2026). Both support more than 2,000 production-ready voices across 100-plus languages and dialects, and both scored top marks on the independent Hume AI voice-quality benchmarks, ahead of the prior Gemini 3.1 Flash TTS generation (Unite.AI, 2026).
For a UK SME, the practical split is simple: Flash-Lite TTS is the one worth testing first for everyday content — explainer videos, training modules, dubbed product demos — because it is the cheaper, higher-volume option.
What Can Gemini 3.8 Speech Models Actually Control?
The headline change is not the voices themselves — it is the delivery. Previous-generation AI narration tools mostly let you pick a voice and hope. Gemini 3.8 lets you direct the performance line by line: tone, pacing, dialect, whispers, laughs, and sighs can all be scripted, and a native two-speaker mode stages multi-turn conversations from a single script while keeping turn-taking natural (Android Authority, 2026).
That matters because the real cost of AI narration has rarely been the recording itself — it has been the re-recording. A voiceover that stresses the wrong word in a client's company name, or rushes a price point in a training video, sends the whole clip back for a redo. Fine-grained delivery control is Google's attempt to close that gap, and it is the detail worth testing before assuming AI narration saves any time at all.
Is Gemini 3.8 TTS Available to UK Businesses?
Partially. UK businesses can use both models today through the Gemini API and Google AI Studio — voice design from a text prompt, the full 2,000-plus voice library, stage-direction controls, and multi-speaker staging are all available now.
What is not available in the UK is voice replication — cloning a real person's voice from a 30-second sample. Google restricts this feature in the UK, the EEA, Switzerland, India, Texas, and Illinois, most likely because of local biometric-data and consent rules (Google, 2026; Android Authority, 2026). In practice, this rules out one popular use case — recreating a specific presenter's voice at scale — but leaves everything else open. A UK SME can still design a bespoke brand voice from scratch, or license one of the 2,000+ stock voices, with SynthID watermarking and C2PA credentials applied automatically to every clip either way.
Businesses handling any EU-facing content should also note that voice-cloning restrictions sit inside a wider compliance picture — see our guide on EU AI Act compliance if AI-generated audio touches EU markets.
How Much Does Gemini 3.8 Flash-Lite TTS Cost?
Gemini 3.8 Flash-Lite TTS is billed per token through the Gemini API: $0.50 per million input (text) tokens and $6.00 per million output (audio) tokens, an introductory rate that holds until 31 December 2026. From 1 January 2027, output pricing doubles to $12.00 per million tokens (Google AI for Developers, 2026). Audio output consumes roughly 25 tokens per second, so a five-minute narrated video runs to a few thousand output tokens — a fraction of a penny at current rates, before accounting for the editing time saved or lost.
Flash TTS, the creative-direction model, is priced higher — $9.00 per million audio output tokens until the same date, reflecting its heavier use in longer-form, higher-production work such as audiobooks (Google AI for Developers, 2026).
Does Controllable AI Narration Actually Save UK SMEs Time?
This is the question that matters more than the price. AI adoption among UK businesses reached 35% by June 2026, up from around 12% in late 2023, with text generation the most common use case so far (ONS, 2026). Voice is the next obvious layer — but only if it removes work rather than relocating it.
The test is simple: does scripting stage directions and vocal bursts take less time than re-recording a mispronounced line would have taken anyway? For short, low-stakes content (internal training, drafts, social captions) the answer is usually yes. For client-facing material — a proposal narration, a branded explainer — the safer approach is still a human proofing pass before publishing, because names, prices, and technical terms are exactly where automated pronunciation still slips.
This launch also lands inside a broader pattern of fast-moving model releases this year, from Grok 4.5's coding-focused launch to Moonshot's open-weight Kimi K3 release. The common thread for SMEs is the same each time: a new capability is not worth adopting until it has been tested against a real, repeatable task — not a demo. Slotting Gemini 3.8 into an existing content pipeline, rather than bolting it on as a one-off, is where workflow automation for small business tends to pay off.
Is Gemini 3.8 Speech Right for Your UK SME?
Gemini 3.8 is worth testing if any of the following apply:
- You produce recurring narrated content (training videos, product explainers, dubbed demos) where a small time saving compounds.
- Your current voiceover process loses time to re-recording mispronunciations or flat delivery, not to voice selection.
- You do not need to clone a specific person's voice — UK availability covers designed and stock voices only.
- You have, or can access, basic API or Google AI Studio integration.
It is probably not worth the switch yet if your narration volume is low, your content is highly regulated or client-sensitive with no proofing capacity, or your main need is cloning an existing presenter's voice, which is unavailable in the UK. If you are unsure which category you fall into, an AI Readiness Audit is the fastest way to find out before spending API credits on trial and error.
FAQ
What is Gemini 3.8 Flash TTS?
Gemini 3.8 Flash TTS is Google DeepMind's text-to-speech model for expressive, produced audio — voice design, stage directions, and multi-speaker scenes across more than 100 languages. Google launched it on 23 September 2026 alongside the lighter Flash-Lite TTS variant (Google, 2026).
What is the difference between Flash TTS and Flash-Lite TTS?
Flash TTS targets deep creative direction — character voices for games, audiobooks, and podcasts. Flash-Lite TTS targets high-volume, cost-efficient output such as dubbing, voice agents, and bulk narration, at a lower audio output price (Unite.AI, 2026).
Can a UK business clone a real person's voice with Gemini 3.8?
Not through Google AI Studio. Voice replication is unavailable in the UK, the EEA, Switzerland, India, Texas, and Illinois, most likely due to biometric and consent regulation. UK SMEs can still design original voices from a text prompt or choose from more than 2,000 stock voices (Google, 2026; Android Authority, 2026).
How much does Gemini 3.8 Flash-Lite TTS cost?
Through the Gemini API, Flash-Lite TTS is priced at $0.50 per million input tokens and $6.00 per million audio output tokens until 31 December 2026, after which output pricing doubles to $12.00 (Google AI for Developers, 2026).
Is AI-narrated audio from Gemini 3.8 watermarked?
Yes. Every clip carries an inaudible SynthID watermark plus C2PA content credentials, so the audio stays identifiable as AI-generated even after editing or re-encoding (Google, 2026).
Can Gemini 3.8 fix mispronunciation and pacing, not just change the voice?
Yes. Both models follow line-by-line stage directions for tone, pacing, and emphasis, plus scripted vocal bursts such as laughs and sighs, giving direct control over delivery rather than only voice selection (Android Authority, 2026).
Do UK SMEs need coding skills to use Gemini 3.8 TTS?
Basic API or Google AI Studio access is required, which suits a developer or an agency partner more than a non-technical owner working alone. Businesses without in-house development often bring in a workflow automation partner to wire it into existing content processes.
Is Gemini 3.8 suitable for client-facing business content?
It is suitable once a human reviews the output first. The controllable delivery features cut re-recording time, but names, numbers, and technical terms still need a proofing pass before anything reaches a client.
Getting Gemini 3.8 Working for Your UK SME
Gemini 3.8 does not fix AI narration by adding another voice to choose from — it fixes it by letting you direct the performance, so the time you save on recording does not get spent again fixing emphasis and pacing after the fact. For UK SMEs, the near-term opportunity sits in designed and stock voices for recurring content, not in voice cloning, which stays blocked in the UK for now.
AI Advisers works with small and medium-sized businesses from Milton Keynes, helping them test tools like Gemini 3.8 against real content workflows before recommending a rollout. If you are weighing whether controllable AI narration is worth adopting, start with an AI Readiness Audit, or talk to our Milton Keynes AI consultancy team about fitting it into an existing pipeline. For a weekly round-up of launches like this one, subscribe to the AI Brief.
Written by AI Advisers, an AI consultancy for Milton Keynes SMEs that tests new AI tools against real client workflows before recommending them.

