GPT-6 Sol And Luna Debut With 50% Price Cuts, Stronger Caching, But Independent Tests Show Mixed Results

OpenAI has officially released two new models, GPT-6 Sol and GPT-6 Luna, expanding the GPT-6 family alongside the flagship GPT-6 Astra introduced earlier this month. According to the company, the new models deliver near-Astra-level performance in professional work, factuality, coding, and computer use, while cutting API prices by 50% compared to the promotional pricing of the previous GPT-5.6 generation. The company attributes the cost reduction to improvements in caching and inference infrastructure, the savings from which it says are being passed directly to users.
Under the new pricing, GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20 respectively, while GPT-6 Luna is priced at $0.10 per million input tokens and $0.50 per million output tokens, down from $0.20 and $1.20. OpenAI positions Astra as its uncompromising best model for the most demanding projects, while Sol and Luna are designed to make advanced AI practical for higher-volume, everyday workloads.
Benchmark Results and Technical Improvements
On AutomationBench, which tests agents across 47 business tools in sales, marketing, operations, support, finance, and HR, GPT-6 Sol at xhigh effort scored 33.2% at $0.27 per task, outperforming Claude Opus 5 at max effort (26.9%) at roughly 9% of its cost per task. On Agents’ Last Exam, Sol at max effort reached 56.4%, exceeding Claude Opus 5’s highest score at about 60% lower cost. For coding, GPT-6 Sol scored 68.8% on DeepSWE v1.1, within 1.1 percentage points of Claude Fable 5’s best result at approximately 80% lower cost, while Luna’s 66.6% was comparable to Opus 5 and Fable 5 at medium effort at 93–96% lower cost. On OSWorld 2.0, Sol matched Claude Opus 5 (60.5% versus 60.3%) at about one-fifth of the cost.
OpenAI also reports halved factual error rates for Sol relative to its predecessor on its internal evaluation of de-identified conversations where users flagged mistakes, with Luna at higher effort matching GPT-5.6 Sol’s reliability at roughly one-hundredth of the cost. Alignment evaluations show both models improving over their predecessors, including lower rates of deception in deliberately challenging coding scenarios.
Beyond pricing, OpenAI has improved prompt caching to raise cache hit rates by default, offering a 90% discount on cached input-token reads, alongside a monitoring dashboard, diagnostics tools, and controls that let developers adjust reasoning effort and tool availability without invalidating cached context. GitHub reports these changes have cut the share of prompt tokens requiring fresh processing by more than 50% across billions of requests.
GPT-6 Sol and Luna are available starting today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, with Luna also accessible to Free and Go users via the desktop app; both are offered in the API as gpt-6-sol and gpt-6-luna, with a gradual rollout throughout the day.
Independent Analysis Confirms Cost Gains, Mixed Results
Third-party evaluation by Artificial Analysis largely corroborates OpenAI’s efficiency claims while painting a more nuanced picture of the models’ capabilities. Running its Intelligence Index, the firm found that GPT-6 Sol at max effort costs approximately $1.06 per task, roughly half the $1.99 of its predecessor, while Luna drops from $0.18 to $0.07 per task — placing both releases firmly on the cost-efficiency Pareto frontier. Notably, the savings stem entirely from the price cut, as both models consume slightly more output tokens per task than their predecessors.
Performance, however, is a mix of progress and regression. In the Coding Agent Index, Sol improved by 2 points to 57, with gains in Terminal-Bench 4.0 and SWE-Atlas-QnA, while Luna lost 2 points, falling behind in SWE-Atlas-QnA and DeepSWE v1.1. Hallucination rates fell sharply on the AA-Omniscience benchmark — from 92% to 60% for Sol — though this was partly achieved by declining to answer more questions, which reduced Sol’s accuracy by 5 points. The most significant regressions appeared in knowledge-work evaluations: Sol dropped roughly 100 Elo points in GDPval-AA v2.1 and Luna about 75, with manual inspection attributing the declines to shorter deliverables that omit required rubric elements and reduced presentation quality.
Disclaimer
In line with the Trust Project guidelines, please note that the information provided on this page is not intended to be and should not be interpreted as legal, tax, investment, financial, or any other form of advice. It is important to only invest what you can afford to lose and to seek independent financial advice if you have any doubts. For further information, we suggest referring to the terms and conditions as well as the help and support pages provided by the issuer or advertiser. MetaversePost is committed to accurate, unbiased reporting, but market conditions are subject to change without notice.
About The Author
Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.
More articles
Alisa, a dedicated journalist at the MPost, specializes in crypto, AI, investments, and the expansive realm of Web3. With a keen eye for emerging trends and technologies, she delivers comprehensive coverage to inform and engage readers in the ever-evolving landscape of digital finance.



