OpenAI cuts Luna and Terra pricing and adds Sol Fast mode

High-volume API work just got dramatically cheaper on OpenAI’s stack. On July 30, the company cut GPT-5.6 Luna API prices 80% and GPT-5.6 Terra 20%, while leaving flagship GPT-5.6 Sol’s base rates alone and shipping a new Fast mode for customers who will pay for speed.
That is not a marketing refresh. It is a direct rewrite of the unit economics for agents, background jobs, and any workflow that was already routing bulk calls to the cheaper GPT-5.6 tiers.
New Luna and Terra token prices
Per OpenAI’s pricing post, API rates effective July 30 are:
- GPT-5.6 Luna: $0.20 per million input tokens and $1.20 per million output tokens (80% cut)
- GPT-5.6 Terra: $2 per million input tokens and $12 per million output tokens (20% cut)
- GPT-5.6 Sol: unchanged
OpenAI describes Luna as its fastest and most affordable model, and Terra as the balanced model for everyday work. The same cuts flow into how Luna and Terra usage is counted against paid subscriptions in Codex and ChatGPT Work: subscription prices and quota budgets stay put, but Terra and Luna now consume fewer credits. Pricing changes begin rolling out on AWS the same day.
The company also claims Luna delivers performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task, and at nearly nine times the speed. That is OpenAI’s own framing. Treat it as a vendor benchmark until independent cost-per-task runs catch up.
What Sol Fast mode costs and gains
Sol does not get cheaper. It gets a faster lane.
OpenAI is introducing Fast mode in the API, replacing Priority Processing. For GPT-5.6 Sol, Fast mode delivers up to 2.5× the speed of Standard processing at twice the Standard price, with no change in intelligence. The switch is backward compatible: requests already tagged priority automatically use Fast mode.
That split is deliberate. Luna and Terra absorb the volume economics. Sol stays the premium reasoning tip, with an explicit speed premium for teams that will pay 2× to cut latency.
Efficiency claims behind the cuts
OpenAI ties the cuts to a feedback loop it says it is already running in production. Within a human-led process, GPT-5.6 Sol autonomously rewrote and optimized production GPU kernels, designed and ran hundreds of experiments on token generation, and monitored training runs. OpenAI credits the kernel work with cutting end-to-end serving cost by about 20%, and the speculative-decoding experiments with raising token-generation efficiency by more than 15%.
Read that carefully. The lab is claiming its frontier model helped shrink the cost of serving itself, then passed those gains to Luna and Terra customers. If that loop holds, price cuts become a recurring product feature, not a one-off promo. If it does not, this is still a sharp competitive move on the cheap and mid tiers while Sol holds the margin line.
Who wins in the API price war this week
Builders already defaulting to Luna for tool-calling loops and background agents win immediately: the same call now costs one-fifth of yesterday’s Luna bill. Terra buyers get a quieter but real 20% cut on the everyday workhorse.
OpenAI is also pushing the routing story: use Sol where uncertainty is expensive, Luna where the plan is clear and the volume is high. Replit, Notion, Ramp, Blitzy, Cognition, and Dust appear as customer quotes on the announcement page.
The competitive stake is the floor. An 80% Luna cut resets what “cheap enough for agent loops” means this week. Rivals on Haiku-class and Flash-class pricing now have to answer with rates, not blog posts. Watch whether Anthropic, Google, and open-weight hosts match within days, and whether AWS-hosted OpenAI pricing finishes rolling through without lag. The discount hits the invoice before the marketing decks catch up.



