OpenAI’s Ultrafast Mode: 14x Speed for GPT-5.6 Sol
OpenAI launches Ultrafast, a new API tier powered by Cerebras that runs GPT-5.6 Sol at 14x speed and 750 tokens per second.

The update
OpenAI has launched Ultrafast, a new API service tier for its GPT-5.6 Sol model. The tier is powered by a partnership with chipmaker Cerebras and delivers up to 750 output tokens per second, a speed increase of 14x compared to standard processing. The service is currently available in preview to a limited group of customers.
Why it matters
Until now, achieving real-time speed with frontier models required trading off model size or intelligence. Ultrafast aims to change that by delivering high-speed performance on the most capable model. This shift enables new workflows where latency is critical, such as real-time incident response, financial market analysis, and live customer support.
What to watch
OpenAI has not yet announced pricing for the Ultrafast tier. As a preview, availability and capacity may expand over time. The specific model designation, “Sol,” is used for this preview and may change in a full release. The technical architecture relies on Cerebras systems to achieve these performance metrics.
Sources
- OpenAI News — Core product details, speed metrics, and use case scenarios.
- techcrunch.com — Corroboration of speed metrics and partner details.


