OpenAI introduced two new models this week: GPT-6 Sol and GPT-6 Luna, both priced significantly below their predecessors. GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, compared to $4 and $20 for GPT-5.6 Sol. GPT-6 Luna is priced at $0.10 for input and $0.50 for output, down from $0.20 and $1.20 respectively.
The company attributed the price reductions to improved inference efficiency and caching capabilities. OpenAI told VentureBeat that the new GPT-6 rates are permanent, while GPT-5.6 pricing is promotional. The premium GPT-6 Astra tier remains the most expensive option at $10 for input and $50 for output tokens.
Shifting Competitive Landscape
OpenAI's pricing move coincided with Anthropic's announcement of its Claude Opus 5.5 model on September 22, priced at $4 per million input tokens and $20 for output—a 20% reduction compared to Opus 5. Anthropic claimed users will spend approximately 40% less overall due to lower token consumption.
The competitive focus has shifted from benchmark performance to practical economics. According to OpenAI, GPT-6 Sol achieved a 32% score on AutomationBench at a cost of $0.34 per completed task, while Claude Opus 5.5 scored 40% but cost $1.28 per task. In the Agents' exam, OpenAI stated that Sol matched Opus 5's maximum performance at 60% lower cost per task.
Emerging Competition
Cheaper alternatives continue to gain traction. OpenRouter data shows DeepSeek doubled its token share on the platform from 9% in January to 18% in June. Chinese models collectively surpassed US models in token share in early June. Xiaomi released its MiMo models under an MIT license, allowing self-hosting as an alternative to per-token pricing.
According to the Ramp AI Index, based on spending data from over 70,000 companies, Anthropic adoption stands at 43.8% while OpenAI is at 39.8%, indicating how readily corporate spending shifts between providers.
The Paradox of Cheaper Tokens
Lower per-token costs may not reduce overall AI spending. Gartner predicted in August that inference costs per agentic workflow will rise more than fivefold through 2028, describing this as the inference paradox: cheaper unit economics can increase overall spending as workflows grow more complex.
Research from the University of Oxford found that while quality-adjusted inference prices fell significantly—more than seven times faster than conventional measures show—measured per completed task, buyer costs stopped falling as reasoning models consumed more tokens. Cheaper tokens may make it economical to deploy AI more extensively rather than reduce total expenditure.


