Editorial illustration for SpaceXAI's Grok Voice Transcribe 2.0 Claims Double Accuracy at USD 0.10 per Hour
Grok Voice Transcribe 2.0: Double Accuracy at $0.10/hr
SpaceXAI put a price tag and a number on its latest transcription model this week: $0.10 per hour of audio, with accuracy the company says is double that of Grok Voice Transcribe 1.0. The model, listed under the ID grok-voice-transcribe-2.0, is live now as a hosted API, running in both batch and real-time streaming modes. There's no open-weights release, so anyone who wants it has to go through SpaceXAI's own service.
The target isn't clean studio audio. SpaceXAI built this version for the audio that usually breaks transcription systems, garbled phone lines, overlapping speakers, regional accents, and spoken-out credentials like phone numbers and email addresses. That focus traces back to where the underlying model already lives: inside Grok Voice, which SpaceXAI says handles tens of thousands of customer support calls daily, transcribes millions of hours of video narration, and powers the Grok assistant in Tesla cars.
The training data comes from that same real-world noise rather than curated recordings, followed by a round of post-training refinement. SpaceXAI is backing the accuracy claim with a specific leaderboard result, which the company lays out next.
SpaceXAI has released Grok Voice Transcribe 2.0, its newest speech-to-text (STT) model. The development team claims it to be twice as accurate as Grok Voice Transcribe 1.0 at the same price.
Why this matters
For developers building call center tools, voice agents, or transcription pipelines, the pitch here is straightforward: same $0.10/hour price, twice the accuracy, and a top spot among 32 models on Artificial Analysis's streaming leaderboard. If that benchmark holds up outside SpaceXAI's own testing, it's a real signal for anyone weighing STT vendors on noisy, real-world audio rather than clean lab samples. The focus on accents, crosstalk, and spoken credentials suggests SpaceXAI is chasing the messy use cases that actually break most transcription models, phone support lines, multilingual customer bases, dictated ID numbers. That's a sensible bet if you're trying to win enterprise contracts where "works great in demos" isn't good enough.
Still, one leaderboard rank against 31 competitors on an 8-hour benchmark set is a thin data point. We'd want to see how grok-voice-transcribe-2.0 performs on your own audio, especially heavily accented or degraded calls, before swapping out a production pipeline. Founders should also watch whether SpaceXAI publishes fuller benchmark details or leaves this as a marketing claim. The API is live now under that model ID, so testing it costs little more than curiosity.
Common Questions Answered
What is the pricing and accuracy improvement of Grok Voice Transcribe 2.0 compared to version 1.0?
Grok Voice Transcribe 2.0 maintains the same pricing of $0.10 per hour of audio while delivering double the accuracy of its predecessor. This means users get significantly improved transcription quality without any additional cost increase.
What deployment modes does Grok Voice Transcribe 2.0 support?
Grok Voice Transcribe 2.0 is available as a hosted API that runs in both batch and real-time streaming modes. This flexibility allows developers to choose the processing method that best fits their application requirements.
What types of audio challenges is Grok Voice Transcribe 2.0 designed to handle?
Grok Voice Transcribe 2.0 is specifically built to handle noisy, real-world audio rather than clean studio recordings, with particular focus on accents, crosstalk, and spoken credentials. This makes it well-suited for call center tools, voice agents, and transcription pipelines that encounter challenging audio conditions.
Is Grok Voice Transcribe 2.0 available as an open-weights model?
No, Grok Voice Transcribe 2.0 is not available as an open-weights release. Users who want to use this model must access it through SpaceXAI's own hosted API service.
How does Grok Voice Transcribe 2.0 rank among other speech-to-text models?
According to Artificial Analysis's streaming leaderboard, Grok Voice Transcribe 2.0 ranks among the top performers, holding a position among 32 models evaluated for speech-to-text capabilities. This benchmark position demonstrates its competitive standing in the STT market.
Further Reading
- Introducing Grok Voice Transcribe 2.0 - SpaceXAI - SpaceXAI
- SpaceXAI launches a $0.20-an-hour transcriber that tops streaming accuracy - RuntimeWire
- SpaceXAI ships Grok Voice Transcribe 2.0 as Loom transcribes every video - RuntimeWire
- Grok lance Voice Transcribe 2.0 avec une précision améliorée - Investing.com
- Release Notes | SpaceXAI Docs - Grok API Documentation - SpaceXAI Docs