Two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, turn voice generation into a creative studio for developers and creators. Flash TTS enables building bespoke character voices from natural language prompts, with granular, line-by-line direction of pacing, emotion, dialect shifts, backchanneling and nonverbal cues; it supports native two-speaker scene staging and long-form generation with minimal speaker drift. Flash-Lite is optimized for high-volume, cost-efficient dubbing, audio content creation and scalable voice agents. Both models run across Google AI Studio, Gemini API/Enterprise/Notebook and Google Vids, and expand an available library from 30 starter voices to 2,000+ production-ready voices across 100+ languages and dialects. Voice remixing (fine-tuning timbre, pitch, pace and accent) is noted as coming soon.
Performance claims include top rankings on Hume AI’s Voice Design Benchmark (#1 overall, 71.4; #1 in accent modeling, 60.8) and #1/#2 on Hume’s Overall Quality Index, plus leading blind-preference results in Voice Arena for multiple languages (Japanese, Brazilian Portuguese, Vietnamese, MSA, Mexican Spanish, Hindi). Practical features include voice replication from a 30-second sample with built-in consent verification, SynthID watermarking and C2PA credentials for provenance, plus tools for saving, managing and scaling custom voices for audiobooks, games, podcasts, dubbing and interactive voice agents.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.