OpenAI Previews an Ultrafast Tier: GPT-5.6 Sol at 750 Tokens per Second, Running on Cerebras

On August 13, OpenAI previewed Ultrafast, a new service tier launching first in the API that it says runs GPT-5.6 Sol up to 14x faster than Standard, generating up to 750 output tokens per second on Cerebras hardware. It is a limited preview for a small set of customers, with no published price, no GA date, and no model ID yet.

What the 14x Is Measured Against

The baseline is Standard-mode inference at roughly 53 tokens per second, and OpenAI says the speedup comes without quality degradation. It also offers cross-vendor comparisons: the accelerated model is described as 11x faster than Fable 5 and 5x faster than Opus 4.8 in Fast mode. Every one of these figures comes from OpenAI and Cerebras themselves, with no independent replication, and OpenAI states that its latency and cost estimates come from simulating production behavior offline — accounting for tool calls and token counts — and that real-world results may vary substantially. The safe reading of "up to 14x" is a ceiling under favorable conditions.

The Pitch Is Latency, Not Capability

OpenAI aims the tier at latency-sensitive workloads: financial research, incident response, customer support, voice applications, commerce, and live experimentation, with early testers reportedly including Jane Street, Podium, Basis, and Rogo. It uses the tier internally for incident response, having engineers analyze logs and traces, summarize conversations, and help prepare or validate fixes. What these jobs share is that model capability is already sufficient — the bottleneck is the dead time a human spends waiting, which is where speed buys more than a stronger model would.

How It Relates to Fast Mode and the Cerebras Deal

Ultrafast is distinct from the existing Fast mode: Fast is pay-as-you-go, promises up to roughly 2.5x for Sol, and costs $10 per million input tokens and $60 per million output, against $5 / $30 for Standard Sol. Ultrafast has no price attached, which means its cost-effectiveness cannot be evaluated yet — for now a team can request access and should avoid writing it into any committed delivery schedule. The backdrop is the multiyear agreement OpenAI and Cerebras announced in January to deploy up to 750 megawatts of Cerebras inference systems in stages through 2028; this preview is the first time that capacity shows up as a specific product tier.

via: OpenAI's announcement, "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed", Cerebras engineering blog, 9to5Mac report; verified 2026-08-15