OpenAI opened a limited preview this week of Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, roughly 14 times its standard processing speed, without any change to the model’s underlying intelligence or context window. The tier runs on wafer-scale hardware from Cerebras, and it represents a genuinely different kind of bet than most recent model releases: rather than chasing a higher benchmark score, OpenAI is betting that latency itself, not just raw capability, is now the binding constraint on where AI can actually be deployed.
What actually changed, and what didn’t
Ultrafast is not a new model, it’s a new way of serving an existing one. GPT-5.6 Sol on Ultrafast produces the identical intelligence and output quality as GPT-5.6 Sol Standard; only the rate at which tokens arrive changes. Standard processing runs at roughly 53 tokens per second, so the jump to 750 tokens per second is a genuine, order-of-magnitude shift, not an incremental tweak. Cerebras says Ultrafast runs about 5 times faster than Claude’s own accelerated Fast mode and 11 times faster than Claude Fable 5 on output speed, though these are Cerebras’s own comparisons rather than independently verified figures, and OpenAI has been explicit that its 14x estimate reflects its own service configurations rather than a direct, apples-to-apples comparison against competitors.
The scale of the demonstration is notable on its own terms. Cerebras says GPT-5.6 Sol on Ultrafast completed Humanity’s Last Exam, a 2,500-question graduate-level test spanning chemistry, economics, and literature, in just over 11 hours, with accuracy comparable to Claude Fable 5, which took more than three days of continuous compute to finish the same test. On GDP-Val, a benchmark built around real, paid knowledge work like legal briefs and financial models, Cerebras claims a 5.6x end-to-end speedup with no measurable drop in output quality.
Why speed unlocks work that intelligence alone can’t
OpenAI’s own framing is direct: until now, getting real-time speed typically meant settling for a smaller, less capable model, forcing a straight tradeoff between speed and intelligence. Ultrafast is meant to remove that tradeoff for a specific category of work where waiting simply isn’t an option. OpenAI names incident response, fraud and market analysis, live customer support, and e-commerce as target use cases, and reports its own internal teams are already using Ultrafast for incident response, where having logs, code changes, and reports analyzed while an outage is still actively happening changes what’s possible in the moment, rather than after the fact.
Early enterprise testers described the shift as qualitative rather than incremental. One tester said the speed completely changes the call experience for complex work, and another described it as unlocking synchronous experiences for users that were previously limited by intelligence, not by what the model could do but by how long it took to say it. That framing matters for anyone building agentic products: an agent that reasons well but responds slowly is a poor fit for real-time interaction, regardless of how capable its underlying reasoning actually is.
What’s genuinely still unclear
Ultrafast remains waitlist-only, available to a small, select group of API customers, with no confirmed general availability date and no published pricing. OpenAI’s existing Fast mode, a separate, already-available tier offering roughly 2.5 times standard speed, is priced at double standard GPT-5.6 Sol rates, but OpenAI has not said whether Ultrafast’s eventual pricing will follow that same multiplier or land somewhere else entirely. Developers evaluating whether to build around this tier should treat both the access timeline and the eventual cost structure as genuinely unresolved rather than assume either will match Fast mode’s existing terms.
The deal also matters considerably for Cerebras itself. The chipmaker went public in one of this year’s larger listings but has since struggled to convince the market that its wafer-scale architecture bet translates into durable commercial demand, beyond the specific, high-profile OpenAI relationship. Cerebras and OpenAI signed a multi-year agreement in January covering up to 750 megawatts of Cerebras inference capacity through 2028, a deal Reuters reported at over 10 billion dollars, though Cerebras later valued the same agreement at more than 20 billion dollars in its own first-quarter disclosure. Neither company has said how much of that contracted capacity is actually supporting Ultrafast specifically, or how it’s being allocated across OpenAI’s broader product line, meaning Ultrafast’s current limited-preview status may reflect real capacity constraints as much as a deliberate rollout strategy.
Part of a broader shift toward monetizing speed itself
Ultrafast is a third tier layered on top of pricing that already distinguishes between speed levels: OpenAI’s existing Fast mode, available today at roughly 2.5 times standard speed, costs about double the standard rate. That structure mirrors how cloud infrastructure providers like AWS have long priced compute, charging more for the same underlying service delivered faster. Applying that same logic to AI inference is a meaningful shift in how frontier labs think about monetization: rather than differentiating tiers purely by model capability, as has been the norm, speed itself becomes a separate axis customers pay for directly. If demand for real-time AI deployment keeps growing the way OpenAI’s own use cases suggest, this tiered-speed model gives the company a direct way to capture revenue from exactly the kind of latency-sensitive workloads that a slower, cheaper model simply cannot serve, regardless of how capable it is.
See OpenAI’s own announcement for the full technical detail behind the preview.




