
xAI has released Grok Voice Transcribe 2.0, a new speech-to-text model designed for real-world audio such as customer-support calls, conversations, spoken credentials and short voice commands. The model became available on September 18, 2026, through the Grok Voice API, with the same pricing as its predecessor.
xAI says Grok Voice Transcribe 2.0 is twice as accurate as Grok Voice Transcribe 1.0 in its internal evaluations while maintaining the previous model’s price. The company also says the new model ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
The launch expands xAI’s voice stack beyond speech-to-speech interactions with a dedicated transcription model aimed at developers building voice agents, dictation systems, customer-support applications and other audio-based workflows.
Quick Summary
- xAI launched Grok Voice Transcribe 2.0 on September 18, 2026.
- The model is designed for real-world speech-to-text applications, including calls, conversations and voice commands.
- xAI reports that it is twice as accurate as the previous model in its internal evaluations.
- xAI says the model ranks first for accuracy among 32 streaming models on Artificial Analysis.
- It supports 25 languages, multilingual switching, speaker diarization and word-level timestamps.
- Developers can use batch transcription at $0.10/hour or streaming transcription at $0.20/hour.
- Atlassian is using the model in Loom for capturing spoken instructions.
What Grok Voice Transcribe 2.0 Changes?
Grok Voice Transcribe 2.0 is built on the audio foundation model behind Grok Voice. xAI says the model was trained on live, noisy and multilingual audio collected across different environments, with additional post-training intended to improve performance on difficult speech conditions.
The focus is not limited to clean recordings from a single speaker. xAI specifically highlights challenging audio involving unreliable phone connections, multiple speakers, accents, phone numbers, email addresses and short voice commands.
In internal evaluations, xAI tested the model against four datasets drawn from production traffic: customer-support telephony, conversations with Grok, spoken credentials such as account codes and email addresses, and short multilingual voice commands.
According to xAI, Grok Voice Transcribe 2.0 improved over Grok Voice Transcribe 1.0 across all four evaluation categories. On the company’s telephony evaluation, it also performed ahead of every model xAI tested.
Grok Voice Transcribe 2.0 Ranks First on Artificial Analysis
The strongest externally referenced performance claim comes from Artificial Analysis, which evaluates speech-to-text systems using word error rate, latency and pricing.
xAI says Grok Voice Transcribe 2.0 currently ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Artificial Analysis measures streaming transcription using its AA-WER Streaming Index, which combines audio from the AA-AgentTalk, VoxPopuli and Earnings22 datasets.
The benchmark is focused on the percentage of words transcribed incorrectly, meaning lower word error rate indicates better transcription accuracy. The current Artificial Analysis leaderboard evaluates streaming models across roughly eight hours of audio covering different accents, domain-specific language and challenging acoustic conditions.
This distinction is important when interpreting xAI’s accuracy claims. A first-place position on a particular speech-to-text accuracy benchmark does not mean the model is universally the best option for every transcription workload, since factors such as latency, languages, audio conditions, pricing and application requirements can affect model selection.
Multilingual Transcription Gets a Major Upgrade
Multilingual speech is another area where xAI reports a substantial improvement.
Grok Voice Transcribe 2.0 can automatically detect languages and handle language switching during a recording in a single transcription pass. xAI says multilingual accuracy represents its largest improvement over the previous model.
On one internal short-phrase evaluation, the reported word error rate fell from 20.6% with Grok Voice Transcribe 1.0 to 6.8% with version 2.0.
The model is also designed for short voice-assistant commands, where there may be little surrounding context to help determine the speaker’s language. This could be particularly relevant for voice interfaces and embedded assistants that need to process brief spoken instructions.
Features for Voice Agents and Developers
Grok Voice Transcribe 2.0 supports both recorded-audio transcription and real-time streaming.
Developers can use the API for batch transcription of audio files and URLs or stream audio for lower-latency transcription. Existing Speech-to-Text API integrations can receive the accuracy improvement without code changes when moving to the new model.
Key capabilities include:
- Word-level timestamps with confidence scores
- Speaker diarization
- Multichannel transcription for up to eight channels
- Key-term biasing for domain-specific vocabulary
- Automatic formatting of numbers, dates, currencies and contact information
- Filler-word removal
- Smart turn detection for voice-agent applications
- Batch and real-time streaming transcription
Key-term biasing can be useful for applications dealing with specialized terminology. xAI allows developers to provide domain terms such as product names or medical vocabulary so the transcription system can better account for those terms.
The current xAI documentation lists support for 25 languages and provides both REST-based and streaming Speech-to-Text interfaces.
Atlassian Uses Grok Transcription in Loom
The launch also includes an early production use case involving Atlassian’s Loom.
According to xAI, Atlassian found Grok Voice Transcribe 2.0 more accurate than its existing transcription solution and is using the model to transcribe Loom videos.
The resulting workflow connects voice instructions with software development. A user can record an action plan or change request in Loom, have the speech converted into text, and then pass the transcript to Cursor for code changes.
The example illustrates a broader use of speech-to-text models: transcription can act as an interface between spoken instructions and downstream AI tools rather than simply producing a written transcript.
Pricing Remains Unchanged
Grok Voice Transcribe 2.0 keeps the same pricing as Grok Voice Transcribe 1.0.
| Mode | Price |
|---|---|
| Batch transcription | $0.10 per hour |
| Streaming transcription | $0.20 per hour |
xAI says diarization, timestamps and key terms are included at these rates.
The model is currently available through the Speech-to-Text API. xAI says Grok Voice Transcribe 2.0 will soon become the default model, while Grok Voice Transcribe 1.0 will be deprecated in the coming weeks. Developers that need to keep using version 1.0 during the transition can explicitly pin the grok-voice-transcribe-1.0 model identifier.
What the Launch Means for Voice AI?
Grok Voice Transcribe 2.0 gives developers a dedicated transcription layer for applications that need to turn real-world speech into structured text before another AI system processes it.
Its combination of streaming support, speaker identification, multilingual transcription, domain-specific term biasing and smart turn detection is particularly relevant to voice agents and conversational applications.
The model also demonstrates how speech-to-text is becoming an increasingly important component of larger AI workflows. In the Loom example, transcription is not the final output; it becomes structured context that can be passed into another AI development tool.
For developers building voice interfaces, customer-support systems, dictation products or agentic applications, Grok Voice Transcribe 2.0 therefore represents an update focused primarily on transcription accuracy and the reliability of speech as an input to downstream AI systems.
Also Read –
Grok Voice Mode Guide: Multimodal Hands-Free AI (2026)
