hn.today

Grok Voice Transcribe 2.0

x.ai24 points7 comments
Screenshot of Grok Voice Transcribe 2.0

Grok Voice Transcribe 2.0 is a new speech-to-text model optimized for noisy, real-world audio. Built on the Grok Voice audio foundation, it was trained on diverse live multilingual recordings and refined post-training to deliver substantially lower word error rates than its predecessor - claimed to be twice as accurate as Grok Voice Transcribe 1.0 at the same price - and to top a public leaderboard of 32 streaming models. Internal evaluations across telephony customer-support calls, conversational Grok interactions, spoken credentials, and short multilingual voice commands show improvements in every category, with the largest gains in multilingual short phrases (WER falling from 20.6% to 6.8%) and telephony leading competitors in production-like tests.

The model is drop-in compatible with existing Speech-to-Text API integrations and supports both batch and streaming transcription, word-level timestamps, free speaker diarization, multichannel (up to 8) transcription, key-term biasing, automatic language detection and mid-recording language switches, text formatting, filler-word removal, and smart turn detection for voice agents. Pricing remains $0.10 per hour for batch and $0.20 per hour for streaming, and the 2.0 model will become the default in the API as 1.0 is deprecated. Early enterprise adoption includes Atlassian Loom, which switched to Grok for more accurate transcripts that feed downstream AI workflows.

Read on x.ai7 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.