Skip to main content
OpenAI logo with a stylized "Ultrafast" speed graphic, symbolizing GPT's 14x performance boost.

Editorial illustration for OpenAI Launches ‘Ultrafast’ Mode, Boosting GPT Speed by 14x

OpenAI Launches ‘Ultrafast’ Mode, Boosting GPT Speed by 14x

3 min read

OpenAI turned on a new gear for its flagship model this week. The company announced Ultrafast, a mode built for GPT 5.6 Sol that pushes output to 750 tokens per second, about 14 times the pace of standard processing. Tokens are the individual chunks of text a model spits out as it answers a prompt, so the jump translates into responses that land in a fraction of the usual time.

The launch, detailed in a blog post published Thursday, leans on OpenAI's partnership with chipmaker Cerebras to hit those numbers. Anthropic already offers a fast mode for Claude, but OpenAI says its version outpaces that option by a wide margin. The company is pitching Ultrafast at businesses running time-sensitive operations: incident response teams, customer support desks, financial analysts tracking fast-moving markets, and e-commerce platforms that need instant replies.

For now, access is limited. Ultrafast is in preview and only a small set of customers can use it, with OpenAI promising wider rollout as capacity allows. The company frames the release as evidence that speed gains no longer require shrinking a model down. Here's how OpenAI described the shift in its own words.

The company says that Ultrafast can work at 14x the speed of standard processing, delivering up to 750 output tokens — such tokens represent the distinct pieces of text generated by an LLM when it interacts with a human — per second.

Why this matters

For developers building on OpenAI's API, 750 tokens a second changes what's practical to ship. Real-time voice agents, live code completion, interactive tools that used to feel laggy under GPT's usual pace, all of that gets easier to justify once latency stops being the bottleneck. But we'd hold off on the applause until pricing and rate limits show up.

Speed claims from labs tend to arrive months before the fine print on cost per token or context window tradeoffs, and 14x faster is meaningless if it's gated behind enterprise tiers or throttled after a few requests. Anthropic's fast mode for Claude suggests this is becoming a competitive front rather than a one-off feature, which is good news for anyone comparison-shopping models. Still, "more useful work per second" is OpenAI's framing, not an independent benchmark.

Founders weighing GPT 5.6 Sol against alternatives should ask what accuracy or reasoning quality gets traded for that speed, because raw throughput numbers rarely tell the whole story on their own.

LIVE22:19Claude Agents Sabotaged Each Other on Shared Server, Hid Actions From Users