Previewing Ultrafast: GPT-5.6 Sol at Up to 14x the Speed
OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than the Standard tier, powered by a partnership with chip maker Cerebras.
Details
- Up to 14x the speed: Ultrafast delivers up to 750 output tokens per second, compared to the Standard tier — a dramatic latency reduction for the same underlying model
- Powered by Cerebras: The speedup comes from running inference on Cerebras hardware, which OpenAI has used previously for ultra-low-latency inference and is now extending to its most capable model, GPT-5.6 Sol
- Sol only, for now: Ultrafast currently supports GPT-5.6 Sol exclusively; OpenAI hasn’t said whether other models in the GPT-5.6 family will get the same tier
- No traditional speed-for-size tradeoff: OpenAI is positioning Ultrafast as delivering frontier-level intelligence without forcing developers to drop down to a smaller, less capable model just to hit low-latency requirements
- Limited preview access: Availability is currently restricted to select customers; OpenAI says it will expand access “as capacity grows” and is taking sign-ups at openai.com/form/ultrafast/
- Pricing not yet disclosed: OpenAI has not published Ultrafast-tier pricing alongside the preview announcement
How to try it
Ultrafast is aimed at use cases where response time is critical to the product experience: incident response and system-outage analysis, financial market analysis and fraud detection, real-time customer support, e-commerce checkout assistance, and interactive research where rapid back-and-forth matters. Access is currently limited to selected customers during the preview; developers can register interest at openai.com/form/ultrafast/ for notification when broader availability opens. Full details are available on OpenAI’s blog at https://openai.com/index/previewing-ultrafast.