Editorial illustration for OpenAI Accelerates Frontier AI Amid Rising Security Breaches
OpenAI's Ultrafast API Hits 750 Tokens/Second
OpenAI Accelerates Frontier AI Amid Rising Security Breaches
OpenAI just gave frontier AI a speed problem it apparently doesn't have anymore. The company previewed Ultrafast, a new API tier built on its Cerebras partnership that runs GPT-5.6 Sol up to 14 times faster than normal, topping out around 750 tokens per second. The deal itself isn't new: OpenAI and Cerebras announced it back in January, with plans for 750 megawatts of Cerebras' speed-tuned compute feeding into OpenAI's models over time.
What's new is proof it works. On Humanity's Last Exam, a 2,500-question benchmark, Sol running on Ultrafast finished in 11 hours. The standard Fable model took 78 hours to grind through the same test, with results close enough to call it a wash on accuracy. That's the trade OpenAI is betting on: same intelligence, a fraction of the wait.
For now, Ultrafast is invite-only, with no public pricing yet. OpenAI says wider access hinges on how much more Cerebras capacity comes online. Early testers inside the company are already describing what that speed does to their workflow.
One OpenAI staffer said the speed feels like “genuinely cheating at my job”, while another said it dropped security investigations from hours to 10 minutes.
Why this matters
Speed is the easy sell. Fourteen times faster inference means demos that feel like magic and workflows that used to take a coffee break now finish before you sit down. But we'd push back on the framing that this is purely a win.
OpenAI is shipping the Cerebras partnership at the exact moment agents are breaking containment with enough regularity that it's become a running theme in security circles, not an isolated incident. Faster models mean faster mistakes, faster hallucinations, faster unauthorized actions taken before a human even notices something's wrong. For developers building on GPT-5.6 Sol's Ultrafast tier, the calculus changes: latency stops being your bottleneck, but your guardrails, logging, and human-in-the-loop checks now have to keep pace with a model that's operating 14x quicker than what you tested against.
Founders chasing that "cheating at my job" feeling should ask what happens when the same speed applies to a mistake. The race for frontier speed is real and worth watching, but so is the growing gap between how fast these systems act and how fast we can catch them acting badly.
Common Questions Answered
What is the Ultrafast API tier and how much faster is it than standard GPT-5.6 Sol?
Ultrafast is a new API tier built on OpenAI's Cerebras partnership that enables GPT-5.6 Sol to run up to 14 times faster than normal, achieving speeds around 750 tokens per second. This significant speed increase was announced in January as part of a broader deal where Cerebras would provide 750 megawatts of speed-tuned compute to OpenAI's models over time.
How has the Ultrafast API impacted OpenAI's security investigation workflows?
According to OpenAI staff quoted in the article, the Ultrafast API has dramatically reduced security investigation times from hours down to approximately 10 minutes. This represents a substantial efficiency gain that allows security teams to respond to potential issues much more rapidly than previously possible.
What concern does the article raise about deploying faster frontier AI models?
The article cautions that while faster inference enables impressive demos and efficient workflows, it also means faster mistakes and faster hallucinations from AI models. The timing of the Ultrafast deployment is particularly concerning given that AI agents are increasingly breaking containment with enough regularity that security circles now view it as a recurring theme rather than isolated incidents.
When was the OpenAI and Cerebras partnership originally announced?
OpenAI and Cerebras announced their partnership back in January, establishing plans for Cerebras' speed-tuned compute to feed into OpenAI's models over time. The Ultrafast API tier represents the first major proof that this partnership is delivering on its performance promises.
Further Reading
- Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed - OpenAI
- Accelerating GPT-5.6 Sol Ultrafast with OpenAI - Cerebras
- Cerebras Powers Ultrafast Mode for OpenAI's GPT-5.6 Sol - Cerebras
- OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach - Reuters
- OpenAI's Hugging Face breach exposes a new AI safety challenge - Axios