skip to content
The Weighted Average

Agentic Engineering

OpenAI Ultrafast Turns Latency Into a Model Choice

OpenAI's GPT-5.6 Sol Ultrafast reaches 750 output tokens per second in limited preview; teams should price latency before redesigning agents.

A rack of electronic equipment in a dark room
A rack of electronic equipment in a dark room. Photograph by Tyler

OpenAI has put GPT-5.6 Sol into a limited Ultrafast preview that promises up to 750 output tokens per second and up to 14× the speed of Standard processing. The decision is not whether every company should switch models; it is whether a few minutes of waiting is now expensive enough to justify a new inference tier, a new hardware dependency, and a different approval boundary.

That distinction matters because Ultrafast is not a smarter model. It is a faster path to the same frontier model, powered by Cerebras and initially offered to a select group of customers. Teams running incident response, live support, financial research, interactive coding, or any workflow where a human is watching the agent should request access and measure completion time. Teams running asynchronous batch work should wait for pricing and capacity evidence before redesigning their stack.

The latency tier is the product

OpenAI’s Ultrafast announcement frames the release as a way to bring frontier intelligence into work where every second matters. Its examples are revealing: analyze logs during an outage, assess financial signals while conditions change, resolve support issues without breaking a conversation, answer commerce questions before a shopper abandons a cart, and turn an overnight research run into an interactive workday. These are not generic benchmark tasks. They are queues with a human, a customer, or a failing system at the other end.

The service is available first through the API, not as a general ChatGPT toggle. OpenAI says the preview is limited to a select group and will expand as capacity grows. There is no public Ultrafast price or service-level guarantee in the announcement. That omission is not a footnote; it is the commercial fact buyers should carry into every conversation. A speed ceiling without a price is a benchmark, not a business case. TechCrunch’s report on the launch confirms the same boundary: the feature is a preview, access is limited, and expansion depends on capacity.

The underlying model already has a cost ladder. OpenAI’s GPT-5.6 release lists Sol at $5 per million input tokens and $30 per million output tokens. Its price-performance update describes Fast mode as up to 2.5× faster than Standard at twice the price. Ultrafast adds another speed-and-cost choice, not a replacement for the existing menu. For context, OpenAI’s Business pricing page lists a $20-per-user-per-month workspace while Enterprise pricing is custom; neither figure prices this API tier or the implementation around it. Procurement should ask which step of a workflow earns the premium rather than route every token through the fastest lane.

Cerebras is the missing half. Its partnership account says its wafer-scale architecture keeps model weights in 44 GB of SRAM on each wafer-sized chip, reducing the memory movement that slows large-model inference on conventional systems. The speed belongs to a particular hardware path, not to GPT-5.6 Sol as a portable property. The earlier GPT-5.6 analysis treated capability as the contest for the default agent slot; Ultrafast adds latency as a second axis. The Gemini Flash economics analysis made the older cost-per-completed-job case; this release puts time-to-completion beside it.

That makes the release an agent-economics event. Model capability determines whether an agent can complete a task. Latency determines whether the task belongs on the critical path at all. A customer-support assistant that returns a correct answer after the conversation has ended is a different product from one that keeps pace with the customer. An incident agent that summarizes the outage after the postmortem is useful; one that helps an engineer test the next hypothesis while the service is failing can change the loss curve.

The difference is not abstract. A fast model can allow a workflow to run fewer parallel sessions, reduce context switching, and keep a human in one decision loop. But it can also make an expensive or unsafe loop execute more rapidly. The procurement question is therefore “which seconds change the outcome?” not “which model has the largest tokens-per-second number?”

A workday beats a leaderboard

Cerebras supplies a more concrete comparison than “up to 14×.” In its head-to-head evaluation, GPT-5.6 Sol Ultrafast answered 2,500 questions from Humanity’s Last Exam in 11 hours and 11 minutes. Claude Fable 5 took 78 hours and 27 minutes to reach comparable conclusions. Divide the two elapsed times—78.45 ÷ 11.18—and the result is approximately 7.0×. That is the useful derived figure: not a vendor’s speed ceiling, but the amount of wall-clock time the particular evaluation removed.

Ultrafast turns a three-day run into a workday

Elapsed time to answer 2,500 Humanity's Last Exam questions

GPT-5.6 Sol UltrafastClaude Fable 50h20h40h60h80h11.2 h78.5 h
GPT-5.6 Sol UltrafastClaude Fable 50h20h40h60h80h11.2 h78.5 h
Cerebras head-to-head evaluation · July 2026

The comparison is not an independent production benchmark. Cerebras ran GPT-5.6 Sol Ultrafast with Codex at xhigh reasoning on July 10 and ran Claude Fable 5 with Claude Code on July 13–15. The workloads, dates, harnesses, and infrastructure are not identical enough to turn the result into a universal ranking. Still, the operational shape is legible. A research team that waits three days for a result can ask a new question before lunch. An incident team that waits minutes for a diagnosis can test a second hypothesis while the outage is still unfolding.

Cerebras also reports a 5.6× end-to-end speedup on GDP-Val with no quality degradation. “End-to-end” matters more than token rate because the clock includes the path around the model. Retrieval, tool calls, orchestration, and rendering all remain in the workflow. If generation is only one quarter of total elapsed time, a 14× generation improvement cannot make the whole application 14× faster. The company’s 5.6× result is therefore a better upper-bound intuition for a real task than the 750-token headline.

The output ceiling is still large. 750 tokens/second × 3,600 seconds gives 2.7 million output tokens per hour under the stated maximum; the comparable Standard ceiling implied by 14× is roughly 54 tokens per second. Both are derived ceilings, not an SLA. Queueing, batch shape, reasoning mode, context length, tool latency, and rate limits can move the observed number in either direction.

OpenAI’s latency guide says fewer requests, parallel execution, streaming, and smaller models can matter as much as raw generation speed. It also notes that cutting output tokens often improves latency more than cutting input tokens. A fast model cannot rescue an agent that serializes ten unnecessary calls or emits verbose intermediate artifacts no person will read.

OpenAI’s Programmatic Tool Calling documentation turns that principle into an architectural recommendation: use code to coordinate predictable tool calls in parallel, filter large outputs, and return a smaller structured result. That approach can make Ultrafast more valuable because it stops the application from paying frontier-model latency for work ordinary code can do. Speed is leverage only when the harness does not squander it.

The first evaluation should therefore measure time to accepted result, not tokens per second. Record the time from user request to a reviewed artifact, successful incident hypothesis, resolved support case, or merged code change. Run Standard and Ultrafast on the same task set, with the same tools and approval rules. If the model call gets faster but review and tool time do not move, the business case is thinner than the demo.

The queue is still the bottleneck

Ultrafast can remove one queue while exposing three others. The first is capacity. OpenAI explicitly says access begins with a select group and expands as Cerebras capacity grows. A team cannot make a critical workflow depend on a preview tier without a fallback route. The second is price. The announcement gives no per-token or per-request rate, so a buyer cannot yet compare the cost of shaving seconds against the value of shaving them.

The third is orchestration. Agents read context, call tools, wait for services, ask for permission, and retry failures. OpenAI’s Multi-agent guide says parallel delegation helps when tasks are independent, but extra agents can increase token use and may not help when work contends over a shared resource. A fast model can amplify an unbounded loop, so hard boundaries around writes, deployments, financial actions, and security-sensitive changes remain mandatory.

The hardware path introduces a fourth risk. Cerebras is purpose-built for high-throughput inference, but an OpenAI tier still exposes buyers to routing, regional availability, and capacity allocation. Demand a portability drill: if Ultrafast disappears, does the application degrade to Standard or stall? Preserve the fallback before the first production experiment, not after the preview becomes a dependency.

The strongest counterpoint is that the speed claims may not survive independent testing. The preview is selective, “up to” describes a ceiling rather than a distribution, and the HLE comparison uses different harnesses on different dates. The thesis breaks if median time-to-result barely moves, the premium consumes its value, or availability is worse. The evidence that would change the verdict is a public rate card, a latency distribution at several context lengths, an availability target, and independent task evaluations that include tool calls and human review.

There is also a strategic counterpoint for buyers: speed is not always value. OpenAI’s efficiency account says internal work on routing, kernels, speculative decoding, and the agentic harness reduced serving cost by 20% and increased token-generation efficiency by more than 15%. Those improvements can flow into Standard processing too. If the ordinary tier becomes faster and cheaper, Ultrafast’s premium may narrow before a long contract pays back.

The broader market supplies a useful warning. The DeepSeek V4 Pro pricing change makes cost vary by time of day, while Writer’s finished-task economics makes the completed objective the unit of comparison. A buyer can lose money by optimizing the wrong layer: cheap tokens with too many retries, fast tokens with too much review, or a managed task price that hides a human exception queue.

Buy the second, not the slogan

Ultrafast is best understood as a reversible option on interactive work. The IBM and OpenAI enterprise practice is betting that implementation and governance are the enterprise bottleneck. The Pony.ai and Uber fleet plan shows the same principle in physical systems: scale is credible only when operations, ownership, and safety controls are specified. A faster model belongs in that conversation. It earns a production lane only when the surrounding system can use the saved seconds.

  • Incident-response and live-support teams should request access now. Shadow Standard on historical logs or support transcripts with the same retrieval and approval path. The pass condition is lower time to a reviewed action, not a prettier token-rate chart.
  • Batch research and routine code changes should stay on Standard. Use the DeepSeek price clock and Writer’s task economics as cost foils. Switch only when queue delay is a measured source of lost value.
  • Platform teams should build the fallback first. Preserve Standard, log the provider and hardware path, and put a sunset clause around preview access. The cost is a second adapter and a routing test; the payoff is avoiding a production dependency on capacity that is not yet public.
  • Every team should measure the whole loop. Track model time, tool time, queue time, review time, retries, accepted-result rate, and spend per accepted result. If Ultrafast reduces only generation time, fix the harness before buying more speed.

OpenAI is selling a new unit of competition: not intelligence per token, but useful work per second. The first builders to benefit will not be the ones who blindly route everything through the fastest lane. They will be the ones who know exactly which seconds change a decision, what those seconds cost, and how to keep working when the preview ends.

Sources