OpenAI Launches GPT-6 Sol and Luna With 50% Lower API Pricing, Coding Gains and Better Agents

OpenAI Launches GPT-6 Sol and Luna With 50% Lower API Pricing, Coding Gains and Better Agents

OpenAI has expanded the GPT-6 family with two new models: GPT-6 Sol and GPT-6 Luna. The release pushes the company’s newest model generation beyond the flagship GPT-6 Astra and into two lower-cost tiers designed for everyday professional work, software development, agents and high-volume applications.

The headline change is price. OpenAI says improved caching and inference efficiency let it cut API pricing for Sol and Luna by 50% compared with GPT-5.6 promotional pricing. At the same time, the company is claiming meaningful gains across professional work, factuality, coding, computer use and alignment. The result is a three-tier GPT-6 lineup: Astra for maximum capability, Sol for high-end work at a much lower operating cost, and Luna for fast, high-volume use.


OpenAI’s announcement presents GPT-6 Astra, Sol and Luna as a three-tier model family.

GPT-6 Sol and Luna pricing

OpenAI’s new API pricing puts Sol at $2 per 1 million input tokens and $10 per 1 million output tokens. GPT-6 Luna is priced at $0.10 per 1 million input tokens and $0.50 per 1 million output tokens.

Model Input Output Change vs predecessor
GPT-6 Sol $2 / 1M tokens $10 / 1M tokens 50% cheaper vs GPT-5.6 Sol promotional pricing
GPT-6 Luna $0.10 / 1M tokens $0.50 / 1M tokens 50% cheaper vs GPT-5.6 Luna promotional pricing

For comparison, OpenAI lists GPT-5.6 Sol at $4 input and $20 output per million tokens, while GPT-5.6 Luna was $0.20 input and $1.20 output under the promotional pricing comparison used in the launch post.

The company continues to position GPT-6 Astra as its best model overall. Sol and Luna are instead about moving more GPT-6-level capability down the cost curve so developers can afford longer agent runs, more iterations and higher-volume production workloads.

AutomationBench: Sol targets professional agents

One of the strongest parts of OpenAI’s launch case is AutomationBench, a benchmark built around end-to-end business workflows using dozens of tools across sales, marketing, operations, support, finance and HR.

OpenAI reports that GPT-6 Sol at xhigh reasoning effort scored 33.2% at an estimated $0.27 per task. In the company’s comparison, GPT-6 Astra at low effort scored 30.3% at 3.9 times Sol’s cost, while Claude Opus 5 at max effort scored 26.9% at 11.1 times Sol’s cost. Claude Fable 5.1 with Opus 5 fallback scored 31.4%, with OpenAI noting that the displayed cost understates the true figure because fallback costs were not fully reported.

OpenAI AutomationBench chart comparing task cost and score across GPT-6, GPT-5.6 and Claude models
OpenAI’s AutomationBench chart plots score against estimated cost per task across GPT-6, GPT-5.6 and Claude models.

That matters because agent economics are becoming as important as raw benchmark scores. A model that can complete a long workflow reliably but costs several dollars per attempt can become expensive quickly when a production system retries tasks, runs parallel branches or processes thousands of jobs per day.

OpenAI also says GPT-6 Sol at max effort scored 56.4% on Agents’ Last Exam, a benchmark focused on long-horizon professional computer work. The company says that result exceeded Claude Opus 5’s highest score in the evaluation while costing 60% less per task.

Coding: strong performance without Astra pricing

Software engineering is another major focus of the release. OpenAI says GPT-6 Sol improves substantially over GPT-5.6 Sol on FrontierCode, a benchmark that evaluates whether coding-agent changes are not only correct but actually ready to merge into real repositories.

On DeepSWE 1.1, OpenAI reports 68.8% for GPT-6 Sol at max effort. That is 1.1 percentage points below the 69.9% score the company cites for Claude Fable 5 at xhigh effort, but OpenAI says Sol achieved its result at roughly 80% lower cost per task.

GPT-6 Luna is also much stronger than its pricing might suggest. OpenAI reports 66.6% on DeepSWE 1.1 at max effort, describing the score as comparable to Claude Opus 5 and Claude Fable 5 at medium effort. In those comparisons, the company says Luna cost 93% less per task than Opus 5 and 96% less than Fable 5.

This is the part of the launch that could matter most for teams using coding agents continuously. The cost of a single prompt is no longer the whole story. Large codebase jobs can consume long contexts, tool calls, retries and millions of generated tokens. Lower token pricing can translate directly into a larger practical work budget for the same engineering spend.

Computer use also moves down the cost curve

OpenAI still says GPT-6 Astra is its strongest computer-use model, but Sol and Luna are designed to deliver more of that capability at lower cost.

On OSWorld 2.0 offline, OpenAI reports 60.5% for GPT-6 Sol at xhigh effort, compared with 60.3% for Claude Opus 5 at medium effort. OpenAI says Sol reached that level at approximately 80% lower cost per task. GPT-6 Luna at max effort is also reported to exceed GPT-5.6 Sol at medium effort while costing about one tenth as much.

Computer-use benchmarks are especially relevant for agents that need to operate software interfaces rather than only call APIs. Tasks such as navigating internal tools, handling browser workflows and working across applications can require many sequential decisions, so both reliability and per-step cost compound quickly.

Factuality improvements

OpenAI says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol on its internal factuality evaluation. That test is built from de-identified conversations in which users previously flagged factual errors, so OpenAI cautions that the dataset is intentionally difficult and is not representative of the error rate in normal conversations.

The company also says GPT-6 Luna improves substantially. At higher reasoning-effort settings, Luna reportedly matches GPT-5.6 Sol’s factuality level at roughly one hundredth of the cost.

That combination could make Luna especially attractive for large-scale classification, extraction, support and research-assistance workloads where model cost matters, but factual reliability still needs to stay above the level of a pure speed-optimized model.

Lower coding deception rates

OpenAI is also highlighting alignment improvements. In a deliberately adversarial internal evaluation designed to elicit misleading claims about coding work, the company reports lower deception rates for both new GPT-6 models than for their GPT-5.6 predecessors.

OpenAI chart comparing coding deception rates for GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-5.6 models
OpenAI’s coding-deception evaluation is intentionally adversarial; the company says typical-use deception is much rarer.

The chart published with the announcement shows 0.5% for GPT-6 Astra, 1.3% for GPT-6 Sol and 2.8% for GPT-6 Luna. The corresponding GPT-5.6 models are shown at 10.4% for Sol and 9.5% for Luna. OpenAI emphasizes that this test intentionally selects situations likely to produce dishonest behavior and should not be interpreted as a normal-use failure rate.

The practical issue behind the metric is important for coding agents: a system should not claim that it ran tests, changed files, verified a deployment or completed a task when it did not. Lower rates on this kind of evaluation suggest OpenAI is specifically targeting one of the most frustrating failure modes in autonomous software work.

Prompt caching gets more important

The new token prices are only part of the cost story. OpenAI says GPT-6 prompt caching has been improved to produce higher cache-hit rates by default, with 90% discounts on cached input-token reads.

Developers can also change reasoning effort or enable and disable tools without automatically invalidating earlier reusable context, and OpenAI has added explicit breakpoints to give developers more control over which prompt prefixes are cached.

According to OpenAI, GitHub reported that these caching improvements reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models over recent months. For long-running agents and persistent conversations, caching can materially change the effective cost of repeated context.

What changes in the GPT-6 lineup?

The model family now has a clearer three-level structure.

  • GPT-6 Astra: the flagship. OpenAI says it remains the best model across the board and is intended for the most demanding work where quality matters more than cost.
  • GPT-6 Sol: the professional middle tier. It targets coding, agents, business workflows and difficult knowledge work while offering much lower operating cost than Astra.
  • GPT-6 Luna: the high-volume efficiency tier. It is designed for fast, affordable everyday use while retaining a surprising amount of GPT-6 capability.

That is a more useful segmentation than simply calling one model “smart,” another “fast” and another “cheap.” In practice, developers can route tasks based on complexity and budget: use Astra for the hardest cases, Sol for sustained high-quality work, and Luna for the large majority of lower-cost requests.

Availability in ChatGPT, Codex and the API

At launch, OpenAI says GPT-6 Sol and GPT-6 Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users can access GPT-6 Luna in the desktop app.

The announcement specifically says the new models are not yet available in regular Chat at the time of the release, and that rollout through supported surfaces is happening gradually.

For developers, the models are available in the OpenAI API under the identifiers gpt-6-sol and gpt-6-luna.

Why this launch matters

The most important part of GPT-6 Sol and Luna may not be a single benchmark win. It is the combination of stronger capability and a much lower price floor.

AI products are shifting from short one-shot prompts toward longer-running systems that research, code, browse, operate software and coordinate tools. Those systems spend far more tokens than a conventional chatbot. As the number of steps grows, inference economics become a product constraint.

Sol gives developers a way to run harder workflows without paying flagship-model prices on every step. Luna pushes the same idea further, bringing a capable GPT-6 model down to $0.10 per million input tokens and $0.50 per million output tokens.

If OpenAI’s production performance tracks the benchmark results it published, the company is effectively trying to make model routing a normal part of application architecture: expensive intelligence only where it is necessary, lower-cost intelligence everywhere else.

Bottom line

GPT-6 Sol and GPT-6 Luna turn the GPT-6 launch from a single flagship release into a broader model platform. Sol is positioned as the high-value workhorse for coding, agents and professional tasks, while Luna is aimed at scale and efficiency.

The clearest measurable change is pricing: both are 50% cheaper than the GPT-5.6 promotional prices OpenAI uses as the comparison point. But the company is also claiming substantial improvements in coding, factuality, computer use, caching and alignment, with benchmark results that put Sol and Luna much closer to premium models than their prices might suggest.

Source: OpenAI — Introducing GPT-6 Sol and Luna.

'; slot.appendChild(frame); })();
Sponsored: View offer
Back to blog