GPT-6.1 Sol: Ultrafast Changes Codex Economics
GPT-6.1 SolCodexUltrafast
What OpenAI added to GPT-6.1 Sol and Codex
I would not reduce this release to another model speed boost. In its DevDay 2026 overview and ChatGPT Learn reference materials, OpenAI describes several connected changes at once: GPT-6.1 Sol, Fast and Ultrafast speed tiers, Decisions API, and an open-source Codex harness.
The key Ultrafast figure is up to 300 tokens per second and up to 8x faster generation in Codex. The API is said to be up to 6x faster. But this is generation throughput, not a promise that every end-to-end task will be completed that many times faster.
Pricing does not resemble discounting either. At launch, Ultrafast cost six times the standard API rate. That means the frequently repeated claim of a 95% price reduction is not supported by OpenAI’s official materials.
Availability comes with an important caveat. GPT-6.1 Sol is announced for Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, while Ultrafast support for Sol is expected later. The Ultrafast tier itself initially targeted $500 Pro subscribers and eligible Enterprise and Edu workspaces, subject to administrator permissions.
Decisions API addresses a narrower need: it receives a question, a limited set of options and context, then selects an answer. It is an interface for routing, classification and choosing the next step without unnecessary open-ended generation. The open Codex harness, meanwhile, lets teams inspect and modify the layer connecting the model to workflows and integrations.
Why higher speed does not automatically mean cheaper development
Ultrafast genuinely changes the equation where latency matters more than price. Interactive coding, short fix cycles and parallel agent tasks can benefit from faster token delivery even if each request costs more.
The first comparison should be total task completion time, not tokens per second. Planning, tool calls, tests and retries do not disappear. If generation represents only part of the cycle, an advertised 8x improvement can easily become a much smaller real-world gain.
Decisions API looks less dramatic, yet it may be more practical from an engineering perspective. A constrained response set is easier to validate and integrate into a deterministic pipeline. There is less magic, but a clearer contract and more visible failure points.
OpenAI’s strategy is easy to read: segment one model by latency class and charge a premium for urgency. The real test for Ultrafast begins when generation speed meets slow tools, long agent loops and the actual cost of a completed task.