This is the English edition. 한국어판 and 日本語版 are also available.

GPT-6 Sol and Luna Reviewed: Half-Price APIs, Benchmarks and Early X Reactions

2026-09-23 · AI · United States · Zoogom Editorial

#OpenAI#GPT-6 Sol#GPT-6 Luna#GPT-6 Astra#Codex#AI benchmarks#API pricing

Three abstract AI compute cores representing different capability and efficiency tiers

OpenAI announced GPT-6 Sol and GPT-6 Luna on September 22, 2026. GPT-6 Astra remains the company’s most capable model for the hardest end-to-end work. Sol targets complex coding and agent workflows, while Luna is designed for focused, high-volume and cost-sensitive tasks. The defining change is not a claim that every score made a generational leap. It is that Astra-era capabilities are moving into dramatically less expensive models.

The launch data is more complicated than a straightforward upgrade story. OpenAI reports strong Sol results in several agent and computer-use evaluations. Independent testing from Artificial Analysis found that Sol’s broad intelligence and coding scores were generally close to GPT-5.6 Sol, while Luna fell on some coding tests and both models regressed on parts of knowledge-work output. The price reduction is real; a universal quality improvement is not yet demonstrated.

Three takeaways

  1. Sol aims to improve the cost efficiency of complex coding, computer use and multistep agent work. Luna targets extremely low-cost repeated processing.
  2. Both models provide a 1.05-million-token context window and 128,000-token maximum output, but crossing 272K input tokens raises rates for the entire request.
  3. The combined evidence supports a lower cost per completed task more strongly than it supports a broad performance leap. Buyers should test success rate, retries and human review on their own workloads.

Where Astra, Sol and Luna fit

Where Astra, Sol and Luna fit: Model, Best-fit work, Standard API input · output, Independent result, Main caution

API figures are Standard rates per million tokens. The independent task-cost figures are Artificial Analysis measurements for its own evaluation suite, not additional API prices. They should not be treated as the same unit.

Astra remains the choice when failure is expensive and the work demands the strongest available reasoning. Sol is positioned for coding agents, complex research, computer use and workflows that call multiple tools in sequence. Luna is better suited to classification, extraction, short drafts and large-scale candidate generation where failures are inexpensive and validation can be automated.

Routing by risk and complexity is more useful than sending every request to one model. A system might use Luna to produce candidates, validate them with deterministic rules, escalate uncertain cases to Sol and reserve Astra or a human reviewer for consequential decisions.

Tasks routed to three model tiers based on risk, complexity and volume, followed by verification loops

Official specifications: the same long context at very different prices

The official GPT-6 Sol model page describes gpt-6-sol as a model for complex coding and agentic workflows. It supports none, low, medium, high, xhigh and max reasoning effort, with medium as the default. Its context window is 1,050,000 tokens, maximum output is 128,000 tokens and knowledge cutoff is April 20, 2026.

Standard API rates per million tokens are $2 for input, $0.20 for cached input, $2.50 for cache writes and $10 for output. Sol accepts text and images and produces text. It does not directly support audio or video.

The official GPT-6 Luna model page positions gpt-6-luna for focused, high-volume, cost-sensitive work. It has the same reasoning-effort options, 1,050,000-token context window and 128,000-token maximum output. Its knowledge cutoff is May 18, 2026. Rates are $0.10 for input, $0.01 for cached input, $0.125 for cache writes and $0.50 for output per million tokens.

Both models support web search, file search, image generation, Code Interpreter, hosted shell, apply patch, Skills, computer use, MCP, tool search, function calling and structured outputs. Fine-tuning is not supported. A shared tool list does not guarantee equal tool-use reliability. Teams still need to measure whether a model selects the right tool, follows permissions and recovers from failures.

The 272K threshold can change the economics

A 1.05-million-token context window does not mean every token is billed at the base rate. When input exceeds 272K tokens, the input and cache rates for the entire request are multiplied by 2 and output rates by 1.5. The surcharge does not apply only to tokens above the threshold.

If a Sol request crosses 272K by a small amount, its effective input rate moves from $2 to $4 per million tokens and its output rate from $10 to $15. Luna receives the same multipliers. For large document sets, retrieval, staged summaries and cache reuse can therefore cost less than placing everything in one prompt.

Regional processing carries a 10% premium where available. Batch and Flex cost 50% of Standard, while Fast costs twice the applicable rate. EU data residency is available only with Standard processing. Model choice, processing tier, context length and cache-hit rate all belong in the same cost plan.

What OpenAI’s benchmarks show

The OpenAI launch announcement emphasizes agents and computer use.

The factuality result needs its qualifier: the sample consisted of conversations selected to induce errors, not a representative sample of everyday use. It is a useful stress test, not an average error rate or proof of equal performance across languages. Cross-company results also require matching reasoning effort, tools, time limits and cost accounting before they become direct head-to-head comparisons.

Independent testing finds lower costs and mixed quality

Artificial Analysis’s independent review agrees that the cost-efficiency frontier moved, but it does not find a broad capability jump.

On the Artificial Analysis Intelligence Index, the measured cost per task was $1.06 for GPT-6 Sol max, down from $1.99 for GPT-5.6 Sol max. GPT-6 Luna max cost $0.07 per task, down from $0.18 for GPT-5.6 Luna max. Sol’s output per task increased from 29K to 31K tokens, and Luna’s from 41K to 51K. The models did not become cheaper merely by answering less; lower token rates drove much of the savings.

Sol max scored 57 on the Coding Agent Index, up from 55 for its predecessor. Luna max scored 41, down from 43. Sol’s Terminal-Bench 4.0 result rose from 37% to 43%, and SWE-Atlas-QnA from 54% to 58%. Luna’s SWE-Atlas-QnA result fell from 49% to 44%, while DeepSWE v1.1 fell from 66% to 64%.

AutomationBench-AA improved from 60% to 62% for Sol and from 50% to 53% for Luna. The counterevidence came from knowledge work: Sol lost about 100 Elo and Luna about 75 Elo on GDPval-AA v2.1. Luna fell about 45 Elo on AA-Briefcase v1.1, while Sol was broadly flat. Artificial Analysis attributes some regressions to presentation quality and omitted deliverables or scoring criteria rather than pure reasoning failure. That is still a production failure when the task is a report, deck or structured document.

A lower hallucination rate is not automatically higher accuracy

On AA-Omniscience, Sol’s hallucination rate fell from 92% to 60%, while Luna’s fell from 93% to 77%. Those large declines are easy to present as a pure accuracy gain. The fuller result is that Sol’s answer-attempt rate also declined from 99% to 83%, and accuracy fell from 59% to 54%.

Sol became more willing to abstain rather than invent an answer when it did not know. Fewer unsupported claims are valuable, but being wrong less often and answering correctly more often are different outcomes. A production evaluation should track accuracy, abstention, confident errors, retries and human-review rates together.

A visual balance among model capability, completion cost, latency, retries and human review

Reading the X reactions by source type

View OpenAI’s original post on X

The official OpenAI launch post on X says Sol and Luna bring Astra advances into faster, less expensive models and reduce token prices by 50% from GPT-5.6 promotional rates. It is a primary source for product positioning and pricing, not independent quality validation.

View Sam Altman’s original post on X

Sam Altman’s launch post emphasizes improvements in intelligence, alignment, work output, coding and computer use, then focuses on half the price per token and an even lower cost per task. This is the company CEO’s launch claim. Real workload cost still depends on retries and review.

View Tibo Sottiaux’s original post on X

Tibo Sottiaux’s post explains the development direction as using stronger models to make capability more efficient and widely accessible. It is useful as an OpenAI employee’s strategy explanation, but it should not be labeled an independent user review.

View Artificial Analysis’s original post on X

The Artificial Analysis launch thread offers the key counterweight: cost efficiency improved, but broad intelligence and coding indexes are largely similar to GPT-5.6, with gains and regressions depending on the evaluation.

View Aravind Srinivas’s original post on X

Perplexity CEO Aravind Srinivas says Sol reached an Opus 5 level at one-fifth the price in Perplexity’s WANDR evaluation. His description of Sol as a Light-effort orchestrator and Astra as a High-effort orchestrator is a concrete product-routing example. WANDR remains a Perplexity evaluation, not a general independent benchmark.

Individual launch-day reactions were mixed. Some users described Luna as a substantial improvement over its predecessor; others found Sol underwhelming but argued that the lower price offset weaknesses. These short impressions often omit prompts, reasoning effort, tools and account conditions, so they are useful signals for what to retest—not population-level evidence.

What U.S. users should verify

At launch, Sol and Luna are available in ChatGPT Work, Codex and the API. Plus, Pro, Business, Enterprise and Edu users receive access through Work and Codex. Free and Go users can access Luna in the desktop app. The models are not yet in general Chat, and staged rollout can make availability differ by account.

ChatGPT subscriptions and API billing are separate. A Plus or Pro plan does not include API tokens, and Codex allowances cannot be calculated solely from the API rate card. U.S. businesses should also distinguish Standard, Batch, Flex and Fast processing, and confirm whether organizational policy permits each model and tool.

An English benchmark does not guarantee equal performance on every U.S. business workflow. Teams should test legal and policy language, spreadsheets, long-form reports, internal terminology, accessibility requirements and structured deliverables. Luna’s low rate is most valuable when validation is cheap and errors can be detected automatically.

Which model should you choose?

Compare the cost of a completed result, not only token prices. Model calls + failed runs + retries + human review + latency can make the cheapest model more expensive. The reverse is also true: a validated high-volume pipeline can make Luna’s pricing materially important.

Bottom line

GPT-6 Sol and Luna do not replace Astra. They divide the same generation across different price, speed and risk profiles. Sol reduces the task cost of complex agent work, while Luna pushes repeated processing to a much lower price tier.

OpenAI’s benchmarks make a strong case for Sol in agents and computer use. Independent testing also finds regressions in some knowledge-work tasks and Luna coding results. The hallucination improvement includes more abstention and lower measured accuracy, so it cannot be summarized as “more correct.” The safest adoption pattern is workload routing plus an internal evaluation, not an immediate all-at-once migration.

Sources and use notice

OpenAI, ChatGPT, GPT, Codex and related marks belong to their respective owners. This independent article is not sponsored or endorsed by OpenAI. It separates official specifications and vendor benchmarks from Artificial Analysis measurements, company commentary and early user reactions. It does not reproduce official charts, product screens, X screenshots or third-party photography.

Source: OpenAI · Includes original screenshots or graphics