For three years, ChatGPT’s voice mode felt like an insufferable phone call with someone driving through the wilderness with one bar of internet: you said your bit, then waited while the machine thought, transcribed, answered, and read it back. OpenAI’s new GPT-Live model now kills that lag, by running voice natively rather than bolting speech onto a text pipeline, and the result is responses in the 300-millisecond range with actual emotional nuance. It’s close enough to human conversational rhythm that the conversation-killing pause you’d normally brace for is not there. So rejoice, I guess.
That reorganises what the thing is now for; that half-second delay keeps you aware you are talking to software while removing it makes it start feeling more like conversation. This is, of course, the whole point, but it’s also exactly what should scare the shit out of people. These tools have already long been weaponized by phone scammers, and this latency reduction is only going to make things worse.
What changed under the bonnet
The old voice mode was three separate systems bolted together: speech-to-text to hear you, a language model to think, and text-to-speech to answer. Every hop added delay, and by one 2024 analysis the stacked lag before the first word of a reply could reach roughly 1,700 milliseconds. GPT-Live collapses that stack. It processes audio natively in a single model, and it is full-duplex, meaning it can listen and speak at the same time rather than waiting politely for you to finish. That is why you can interrupt it mid-sentence and it adjusts, the way a person does.
The numbers are worth stating precisely, because the marketing rounds them down. OpenAI is targeting 300 milliseconds, with a reported median time-to-first-audio in the 300 to 600 millisecond band for its core model. So it is not always sub-300, but even the upper end is a different category of experience from the old one-to-two-second wait. GPT-Live-1 becomes the default for ChatGPT Voice on Go, Plus and Pro; GPT-Live-1 mini covers the free tier. It rolled out on 8 July, so if your voice mode has felt suddenly less stilted lately, this is why.
The bit that should give you pause
Here is the uncomfortable flip side, and it is not hypothetical. The single thing that has always given a voice scam away is timing: the unnatural pause, the slightly-off cadence, the beat too long before an answer. Sub-second, emotionally-inflected, interruptible voice AI erases exactly those tells. The same upgrade that makes talking to ChatGPT pleasant also makes an automated impersonation harder to catch, and voice-cloning fraud was already a growth industry before this.
None of that is a reason to dismiss the technology, which is genuinely impressive and genuinely useful for translation, accessibility and hands-free work. But it is a reason to update your instincts. The old advice, trust your ear, the robotic pause will give it away, is expiring fast. As real-time voice models spread, verification has to move from how a voice sounds to what the person actually knows, and it is worth telling the less tech-savvy people in your life the same thing before they get a very convincing call.
Related on Top Tool Stack: the wave of new frontier models this autumn.