On August 13, 2026, Cerebras announced it is running OpenAI’s most capable model, GPT-5.6 Sol, at a blistering 750 output tokens per second on its wafer-scale chips. The new OpenAI service tier, called Ultrafast, delivers up to 14× faster inference than Standard processing — a speed no GPU cloud has publicly matched. This removes the fundamental tradeoff between intelligence and latency, making frontier-level reasoning viable for real-time products.

What Happened

Cerebras, the maker of wafer-scale processors, has partnered with OpenAI to power GPT-5.6 Sol on a new Ultrafast tier available through the OpenAI API. The headline number is 750 output tokens per second — fast enough to stream complex reasoning as it happens. But the more revealing comparisons come from Cerebras’ head-to-head benchmarks: against speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast runs 11× faster than Claude Fable 5 and 5× faster than Opus 4.8 on Fast mode.

The companies also stress-tested the tier on Humanity’s Last Exam, a 2,500-question PhD-level benchmark, and achieved results that suggest no accuracy loss at speed. The Ultrafast tier launches as a limited preview for select customers, with access expanding as capacity grows. Both Cerebras and OpenAI frame this as eliminating the compromise between a model smart enough for high-stakes work and one fast enough to use while the work is still happening.

This is not just a speed bump — it’s a step change in what applications can be built with frontier models. Real-time coding assistants, live document editing, interactive agents, and financial analysis that previously required batch processing can now run interactively.

Read the full announcement →

My Take

For over a year, the narrative in AI has shifted from “how smart can we make models” to “how fast can we run them for useful work.” This announcement answers that question decisively. Cerebras has shown that wafer-scale architecture isn’t just a curiosity — it’s a practical advantage for inference, not just training. The 14× speedup over standard GPU-based inference transforms GPT-5.6 Sol from a tool you wait on into one you converse with.

Developers should pay attention: the Ultrafast tier changes the design constraints for agentic workflows. If you can get a frontier model response in under a hundred milliseconds (accounting for network latency), whole new categories of real-time decision-making open up. The limited preview means early adopters will have a competitive window. If you build products that depend on model latency, now is the time to get on the waitlist.

What to Watch

  • Pricing and availability: Ultrafast will likely command a premium, but if throughput scales enough, the cost per token could still be lower than standard for high-volume users. Watch for public pricing tiers.
  • Competitor response: NVIDIA, AMD, and cloud providers will need to match this speed or risk losing high-value inference workloads to Cerebras. Expect announcements in the next quarter.
  • Application paradigm shift: Real-time document drafting, live code review, and interactive agent loops will become standard. Products that previously used smaller, faster models for responsiveness may now justify using frontier models instead.