faster than Standard processing
GPT‑5.6 SOLOPENAI API
THE WAIT
JUST SHRANK.
Ultrafast is a new OpenAI API service tier for GPT‑5.6 Sol, powered by Cerebras—built for work where the answer’s value changes by the second.
output tokens generated per second
LIMITED
PREVIEW
select customers now; access expands as capacity grows
“Up to” is essential: these are announcement maxima, not a per-request guarantee. Output throughput is not the same as total task latency.
Enter the speed labA 60-SECOND MENTAL MODEL
Three layers.
Do not collapse them.
Ultrafast changes the serving path—not the identity of the model. Here is the clean way to hold the announcement in your head.
OPENAI DESCRIPTION
GPT‑5.6 SOL
The flagship model in the GPT‑5.6 family. This is the intelligence doing the work.
THE ANNOUNCEMENT
ULTRAFAST
A new speed class launching first in the OpenAI API. This is how the model is served.
PARTNERSHIP
POWERED BY CEREBRAS
Cerebras supplies the inference technology behind this service tier.
FEEL THE RATIO
Start the same work.
Watch the window open.
Standard is fixed at 1×. You control the comparison factor, capped at OpenAI’s claim of up to 14×. “Work units” and timings below are illustrative—not measured tokens, latency, or production performance.
Move the control to compare the shape of the time gap. It does not predict your request.
STANDARD
ULTRAFAST
Speed matters when the world keeps moving while a response is being made.
In this normalized run, the faster lane finishes while the Standard lane still has work in flight.
ONE SECOND, UNPACKED
What does up to 750 measure?
OpenAI says Ultrafast generates up to 750 output tokens per second. That is output throughput: the rate generated after output begins.
It is
A rate for generated output tokens. Tokens are pieces of text—not a fixed number of words.
It is not
A promise that every response arrives in 1 second, nor a complete measure of time-to-first-token or end-to-end task time.
Keep the qualifier
Up to travels with the number. Workload, configuration, and serving conditions can change observed performance.
THE DEEPER IDEA
More useful work
per second.
OpenAI’s framing is not “speed for its own sake.” The point is to keep frontier intelligence inside a live decision loop—without reaching for a smaller or more specialized model simply to get a real-time response.
FIVE CLOCKS ARE TICKING
Can you spot why speed changes the work?
Dispatch each scenario to the time pressure that makes a faster frontier model useful. Every scenario comes directly from OpenAI’s announcement.
INCIDENT RESPONSE + RELIABILITY
A critical system is failing right now.
Logs, code changes, traces, and engineer reports are arriving while the outage unfolds.
You found the real pattern.
In every case, the environment—or the person—keeps moving. Throughput matters because stale intelligence loses value.
EARLY-CUSTOMER SIGNALS
Four views from
inside the preview.
These are attributed customer statements published by OpenAI—not independent benchmarks. Together they point to a shift from waiting on work to staying inside it.
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”
“For us the Ultrafast has been invaluable in our voice stack. The speed completely changes the call experience for the more complex work.”
“Oftentimes the barrier to truly fast products is not just tokens per second, but also model intelligence, and ultrafast combines both.”
“Speed doesn’t just make the product feel better. It changes what people can realistically use it for. Ultrafast makes complex financial research feel like a real-time interaction.”
REPORTED INTERNAL USE
The loop is already
tightening.
OpenAI says a group of its developers is testing Ultrafast to learn which workflows benefit when frontier intelligence can answer in real time.
INCIDENT RESPONSE
LIVE SYSTEM- 01ALERT FIRESthe system is still changing
- 02READ THE EVIDENCElogs · traces · conversations
- 03TEST THE NEXT CHECKidentify likely cause
- 04PREPARE / VALIDATE A FIXkeep the intelligence of Sol
OpenAI explicitly says engineers remain responsible for judgment and deployment.
LIVE RESEARCH
ITERATIVE- 01SEARCH + QUERYknowledge sources · connected data
- 02GATHER + ORGANIZEkeep evidence in flow
- 03EXAMINE RESULTSlearn while attention is present
- 04ADJUST + RUN AGAINcompress the feedback loop
OpenAI reports seeing this loop tighten; it does not claim every overnight workflow becomes instant.
THE PARTNERSHIP, CLEARLY
SERVICE
≠ SILICON
Ultrafast marks the next step in OpenAI’s partnership with Cerebras for ultra-low-latency inference.
OpenAI
offers the API service tier and publishes the announcement claims.
Cerebras
is the infrastructure partner powering GPT‑5.6 Sol on Ultrafast.
The precise claim
Up to 750 output tokens/s, according to OpenAI, corroborated by Cerebras.
Fast frontier inference—
not general access.
GPT‑5.6 Sol on Ultrafast is available to a select group of customers. OpenAI says access will expand as capacity grows. The service is launching first in the OpenAI API.
Same frontier model. A new API speed class. Cerebras-powered inference. The promise is not merely faster text—it is less time between seeing, deciding, and doing.
TRACE THE CLAIMS
Source notes
& reading key.
This explainer is independently authored from the sources below. It is not an OpenAI product page and makes no claim of affiliation.
PRIMARY SOURCE · OPENAI · 13 AUG 2026
“Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed”
Basis for the up-to-14× and up-to-750-output-tokens/s claims, scenarios, internal-use descriptions, customer statements, partnership, and availability.PRIMARY MODEL CONTEXT · OPENAI · 9 JUL 2026
“GPT‑5.6: Frontier intelligence that scales with your ambition”
Basis for describing Sol as the flagship model in the GPT‑5.6 family.SUPPORTING PARTNER SOURCE · CEREBRAS · 13 AUG 2026