On 29 July, xAI announced Grok Voice Think Fast 2.0, its next-generation speech model, and the numbers are aimed squarely at one target: the humans currently answering phones for a living. It talks back in 0.70 seconds and costs $0.08 a minute.
A quick definition, because this is the interesting part. A speech-to-speech model listens to your voice and replies with its own voice directly, without stopping to type everything out as text in between. That missing step is why the old voice assistants felt like talking to a laggy satellite phone. Cut the round trip and the thing can actually hold a conversation.
Fast enough to feel human
The metric to watch is time to first audio, the pause before the model starts speaking. Version 1.0 took 1.25 seconds. Version 2.0 takes 0.70. That sounds trivial written down, but in a live call it is the difference between a natural exchange and that awkward beat where everyone talks over each other.
xAI says it managed this by having the model reason “in parallel” with speaking, thinking and talking at the same time rather than one after the other, while using around 60% fewer reasoning tokens to do it. On Artificial Analysis’ speech-to-speech benchmark it scored 82.9%, up from 75.7% for version 1.0, and ahead of OpenAI’s GPT-Realtime-2.1 at 79.1% and Google’s Gemini 3.1 Flash at 69.5%. On transcription, xAI claims a 1.5 to 2 times improvement over specialist tools like Deepgram Nova 3 and ElevenLabs Scribe v2, across 24 languages.
Standard caveat applies: the benchmark is a third-party one, but the framing around it is the company’s. Believe the direction, verify the size.
Read the marketing carefully
xAI describes the model as “built for voice agents in the real world”, able to hear clearly in noisy conditions and reason through complex workflows. Decode that. “Real world” and “noisy conditions” means a call centre floor, a drive-through, a support line. “Complex workflows” means the scripted back-and-forth that a human agent currently walks you through. At $0.08 a minute, a machine that never takes a break undercuts a payroll comfortably.
None of this is hidden, to be fair to them, but it is worth saying plainly. This is enterprise kit designed to sit on the phone lines, and the business case is labour cost. The people cheering the latency numbers loudest are the ones who sign the wage bills.
The other side of the cheap
Here is the part that actually helps the little guy. The same $0.08 a minute that lets a big firm thin out a call centre also lets a one-person business have a phone line that answers every call, books appointments and speaks two dozen languages, for pennies. Cheap voice AI is a threat to workers and a gift to sole traders at the same time, and pretending it is only one of those is how you get caught out.
One practical note: xAI says the “grok-voice-latest” API alias switches from 1.0 to 2.0 on 5 August, so if you are already building on it, that upgrade arrives whether you ask for it or not. Test your flows before the date, not after.
Did you know: the old voice assistants felt slow largely because they converted your speech to text, thought in text, then converted the answer back to audio, three separate steps that a speech-to-speech model collapses into one.
Related on Top Tool Stack: ChatGPT Work · DeepSeek’s free V4 Flash