Skip to main content
Executives discuss voice AI's future, comparing it to ChatGPT's impact, with a focus on next-gen development.

Editorial illustration for Execs Say Voice AI Hasn’t Reached Its ChatGPT Moment Yet, Point to Next Stage

Execs Say Voice AI Hasn’t Reached Its ChatGPT Moment...

• 3 min read

Investors have put billions into voice AI over the past two years, backing everything from model makers to customer service platforms, meeting notetakers, and dictation tools. New releases claiming to sound human and hold natural conversation show up almost weekly. Full-duplex models, systems that can speak and listen at the same time, were supposed to be the breakthrough that made voice interfaces feel real.

But PolyAI CTO Shawn Wen isn't convinced the category has hit its version of ChatGPT's launch moment. Speaking on stage at the HumanX conference last month, Wen argued that getting models to talk and listen simultaneously was only the first milestone. The harder problem, he said, is reasoning speed: models need to fetch answers fast enough that a conversation doesn't feel like it's lagging behind a real exchange.

Wen also pushed back on the idea that customer service agents need to sound robotic to be trustworthy. He described a gradual trust-building process instead, where callers warm up to an AI agent over the first few exchanges and, if the system performs, stop feeling like they need a human on the line at all.

Every week there is a new model or a tool release that claims to sound human and converse like one. However, in reality, that might not be the case. Enterprise voice AI platform PolyAI’s CTO Shawn Wen thinks that despite the release of full-duplex models — which can speak while listening to you — voice AI doesn’t have its “ChatGPT moment” yet.

Why this matters

Wen's point about turn-by-turn trust building is the part worth sitting with. ChatGPT's breakout moment came from a single demo convincing people a text box could reason. Voice doesn't work that way: users test an agent for two or three exchanges before deciding whether to keep talking to it or bail to a human. That's a slower, more granular trust curve, and it means the "good enough" bar isn't a model benchmark, it's a behavioral one measured in conversation turns.

For founders in this space, that's a warning against chasing a single flashy demo as the proof point. PolyAI and others building enterprise voice agents are effectively competing for incremental trust, not a viral aha moment, which changes how you'd measure progress and what you'd show investors. For researchers, it suggests benchmarks focused on first-turn naturalness may be missing the metric that actually predicts adoption: whether users stick around for turn four.

Billions are flowing into voice AI on the assumption it mirrors the LLM trajectory. If Wen's right, that comparison itself is the thing to question.

LIVE16:27Execs Say Voice AI Hasn’t Reached Its ChatGPT Moment Yet, Point to Next Stage