OpenAI is previewing Ultrafast, a new API service tier designed to run its GPT-5.6 Sol model at higher speeds. According to OpenAI, the tier delivers up to 14 times faster performance, reaching up to 750 output tokens per second. The service is powered by Cerebras.
Why it matters
Output speed is a key factor for latency-sensitive applications built on large language models. A tier offering substantially faster token generation could affect how developers design and deploy applications using GPT-5.6 Sol.
Who should care
Developers and teams building on the OpenAI API who need faster response times from GPT-5.6 Sol should note the availability of this preview tier.