OpenAI’s New Pitch: More Work Per Second

OpenAI has introduced Ultrafast, a preview mode designed to make GPT-5.6 Sol, its latest and most powerful model, run at up to 14 times the speed of standard processing.
The company says Ultrafast can generate as many as 750 output tokens per second. A token is the basic unit a large language model uses to process and produce text; it can be a word, part of a word, or punctuation. In practice, higher output-token speed generally means users see responses appear faster.
OpenAI’s message is aimed squarely at enterprise buyers. In a blog post, the company said that real-time speed has historically required companies to choose a smaller or more specialized model. Ultrafast, it argues, points to a different approach: delivering “more useful work per second” without moving away from its flagship model line.
What Is Available Now

Ultrafast is not yet a broad public release. OpenAI is making the preview available only to a small group of customers and says access will expand as capacity grows. The company has not disclosed pricing, service-level terms, regional availability, detailed infrastructure specifications, or a date for general availability.
The main facts disclosed so far are:
- Model: GPT-5.6 Sol;
- Speed claim: up to 14x faster than standard processing;
- Throughput: up to 750 output tokens per second;
- Status: preview;
- Availability: limited to a small group of customers;
- Infrastructure partner: chipmaker Cerebras.
That last point is notable. Ultrafast is powered by OpenAI’s partnership with Cerebras, underscoring how much modern AI performance depends not only on model design, but also on chips, inference systems, scheduling, and available compute capacity.
Why Enterprises Care About Faster Frontier Models

OpenAI says the mode can support corporate workflows such as incident response, customer service and support, financial market analysis, e-commerce, and related use cases. These are areas where latency can become a business constraint rather than a minor user-experience issue.
In incident response, teams need to interpret alerts and possible next steps quickly. In customer service, delays can affect satisfaction and case resolution. In market analysis, speed matters because information loses value quickly. In e-commerce, faster generation can make product guidance, support, and transaction-related conversations feel more fluid.
Still, faster generation does not automatically solve every enterprise adoption issue. Companies will still need to evaluate governance, security, reliability, integration with internal systems, and cost. OpenAI has not said that Ultrafast changes GPT-5.6 Sol’s accuracy, context handling, or safety behavior, so the clearest reading is that this preview focuses on speed and throughput.
The Competitive Context

OpenAI is not alone in trying to make large models faster. Competitors including Anthropic have introduced accelerated modes; Claude, for example, has a fast mode. The TechCrunch report notes, however, that Claude’s fast mode does not match the speed OpenAI is claiming for Ultrafast.
This shows how AI competition is moving beyond raw model capability. For enterprise customers, the question is increasingly whether a model can perform well inside real production workflows. Benchmarks matter, but so do response time, scale, availability, and deployment economics.
If OpenAI can expand Ultrafast while keeping performance reliable, it could reduce the traditional trade-off between using a smaller model for speed and a larger model for quality. For now, the limited preview means the market will have to wait to see how the 14x speed claim performs in real customer environments. The broader direction is clear: flagship AI models are being pushed toward lower latency, and speed is becoming a central battleground for enterprise AI platforms.
