OpenAI lowers prices and accelerates GPT-5.6 development
OpenAI has reduced prices for its GPT-5.6 models and introduced a new Fast mode for the API, which speeds up request processing. The Luna and Terra plans have become significantly more affordable, while performance has increased compared to competitors.
Crius
OpenAI has announced a price reduction for its GPT-5.6 model lineup and the introduction of a new Fast mode for its API. These changes affect the Luna and Terra pricing plans, which are among the company’s more affordable options. OpenAI also shared comparative data on the performance and cost of Luna versus similar solutions from Google, Anthropic, and other companies.
Price Reductions for GPT-5.6
The API cost for all GPT-5.6 models has been lowered thanks to improved operational efficiency, a result of advancements in the models themselves. The Luna plan, known for its high speed and accessibility, is now 80% cheaper. The Terra plan, designed for everyday tasks, has seen a 20% price drop. The new prices took effect on July 30.
Subscriptions and limits for ChatGPT and Codex services remain unchanged. However, using the Terra and Luna models now consumes fewer credits within these plans.
New Fast Mode for API
The API now features a Fast mode, replacing the previous Priority Processing tier. For the GPT-5.6 Sol model, Fast mode processes requests up to 2.5 times faster than the standard speed at double the cost, with no change in output quality. All existing requests tagged as "priority" have been automatically switched to the new mode.
Performance and Cost Comparison
In terms of performance, Luna is now comparable to models that were considered cutting-edge just a year ago, with a cost of about 6 cents per task and nearly nine times the speed. In the professional benchmark Agents' Last Exam, Luna outperformed Claude Fable 5, while the cost per task was almost 99% lower.
On the Artificial Analysis v4.1 index, GPT-5.6 Luna scores over 51 points at a cost of around $0.05 per task, which matches or surpasses competitors like Claude Opus 5 Low and Gemini 3.6 Flash, both of which are 5–10 times more expensive.
Reasons for Increased Efficiency
The achieved results are attributed to progress in three areas: model architecture, inference systems, and agent integration, which connects models with tools. The GPT-5.6 Sol model contributed to these improvements by autonomously rewriting production cores and conducting numerous experiments to boost token generation efficiency. Core optimization reduced overall model maintenance costs by 20%, while experiments increased token generation efficiency by more than 15%.
