Cerebras announced on August 13, 2026, that it is powering Ultrafast mode, a new OpenAI API service tier for GPT-5.6 Sol. In a limited preview for select OpenAI customers, Ultrafast runs the flagship model at up to 750 output tokens per second, or up to 14 times faster than Standard processing, without changing the model’s intelligence, according to Cerebras.
The company said the tier is meant for work that cannot wait, from incident response to customer support and other latency-sensitive agent workflows. Access is expanding as capacity grows.
Speed without shrinking the model
In its announcement, Cerebras argued that builders have long traded intelligence for latency by choosing smaller models. Ultrafast is positioned as frontier capability at wafer-scale inference speeds. Cerebras reported that GPT-5.6 Sol on Ultrafast completed all 2,500 questions in Humanity’s Last Exam in 11 hours and 11 minutes, versus 78 hours and 27 minutes for Claude Fable 5 in its comparison, with comparable accuracy. On GDP-Val, it reported a 5.6x end-to-end speedup with no quality degradation.
Cerebras attributes the speed to its Wafer-Scale Engine architecture, which keeps model weights on-chip and reduces the memory-bandwidth bottleneck that slows large-model inference on conventional GPU systems.
Decoded Take
Ultrafast turns inference speed into a product tier, not just a hardware footnote. If OpenAI can scale Cerebras capacity, latency becomes a pricing and packaging lever for agents that need to act while events are still unfolding. The near-term constraint is supply. Watch how quickly the preview opens, and whether competitors answer with their own specialty-silicon speed tiers rather than smaller, dumber models.