OpenAI Unveils GPT-5.6 Sol Ultrafast

OpenAI launches GPT-5.6 Sol Ultrafast mode, up to 14x faster, powered by Cerebras, enabling real-time AI for critical business tasks.

Screenshot showing GPT-5.6 Sol Ultrafast and Standard modes building a 3D warehouse simulator from the same prompt.
OpenAI News
Visual TL;DR
GPT-5.6 Sol UltrafastCore
OpenAI launches new mode for GPT-5.6 Sol, up to 14x faster
From the article 2 mentionsGPT-5.6 Sol on Ultrafast mode is currently in a limited preview.
Powered by CerebrasCore
From the article 4 mentionsThis significant speed boost, achieved through advancements in model efficiency and powered by Cerebras hardware, could redefine what's possible in AI-driven products.
750 tokens/secondContext
From the articleThe Ultrafast tier is capable of generating up to 750 output tokens per second.
OpenAI API firstContext
From the articleThe service launches first via the OpenAI API.
Eliminates speed-intelligence trade-offEffect
aims to eliminate the historical trade-off between speed and model intelligence
Real-time AIEffect
enabling real-time AI for critical business tasks where every second counts
From the article 3 mentionsFinancial Research: Processing market signals and transactions in real-time to identify opportunities or risks.
Critical workflowsOutcome
targeting business applications like incident response for rapid analysis
Redefines AI productsOutcome
From the article 2 mentionsThis significant speed boost, achieved through advancements in model efficiency and powered by Cerebras hardware, could redefine what's possible in AI-driven products.
Contents(3)

OpenAI is rolling out a new speed tier for its flagship AI model, GPT-5.6 Sol, dubbed "Ultrafast mode." This development promises to deliver up to 14 times the processing speed of the current standard offering, aiming to make advanced AI practical for time-sensitive applications. The service launches first via the OpenAI API.

The Ultrafast tier is capable of generating up to 750 output tokens per second. This significant speed boost, achieved through advancements in model efficiency and powered by Cerebras hardware, could redefine what's possible in AI-driven products. Historically, achieving such speeds often meant compromising on model intelligence, forcing developers to choose between faster, less capable models or slower, more powerful ones. Ultrafast aims to eliminate that trade-off.

Real-Time Intelligence for Critical Workflows

OpenAI is targeting business applications where every second counts. Early use cases shared by the company include:

  • Incident Response: Rapidly analyzing logs, code changes, and reports to pinpoint and address system failures as they happen.
  • Financial Research: Processing market signals and transactions in real-time to identify opportunities or risks.
  • Customer Support: Resolving complex customer queries during live conversations without delays.
  • Commerce: Providing instant product information, recommendations, and checkout assistance to prevent abandoned carts.
  • Live Research: Transforming lengthy overnight experiments into interactive, iterative sessions.

These scenarios demonstrate a shift towards AI that can keep pace with human interaction and dynamic environments, rather than requiring users to wait for analysis.

Customer Validation and Cerebras Partnership

Initial customer feedback highlights the tangible benefits of the speed increase. John Crepezzi from Jane Street noted that the speed from Cerebras enables new ways of working with AI models, making it more practical for focused development. Similarly, Courtland Lykins at Podium found Ultrafast invaluable for enhancing voice AI call experiences, especially for complex tasks. Mitch Troyanovsky of Basis pointed out that Ultrafast combines model intelligence with speed, overcoming previous limitations for real-time user experiences. Alex Wang from Rogo emphasized that the speed makes complex financial research feel like a live interaction, expanding practical use cases.

The partnership with Cerebras is central to this advancement. Cerebras has been instrumental in providing the ultra-low-latency inference capabilities that underpin Ultrafast mode. This collaboration underscores a broader trend of specialized hardware accelerating AI model performance.

Internal Use and Future Expansion

OpenAI is also using Ultrafast internally. A team is employing it for real-time incident response, quickly synthesizing alerts and logs to aid engineers in diagnostics and remediation. For research, iterative experimentation loops are being tightened, allowing for multiple hypothesis tests within a single workday, a significant acceleration from previous overnight batch processes.

GPT-5.6 Sol on Ultrafast mode is currently in a limited preview. OpenAI plans to expand access as capacity grows, inviting interested businesses to sign up for notifications. This move signals OpenAI's continued focus on optimizing not just model capabilities but also their efficiency and deployment speed for broader enterprise adoption.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.