The Agentic Post
Breaking
Digital Twins and Physical AI  Â·  Humanoid Robots in Manufacturing  Â·  AI Data Centers and Water Usage  Â·  The AI Chip Supply Chain, Explained  Â·  What Is Fine-Tuning? A Plain Explainer  Â·  Meta and Sierra Want to Give AI Agents a Front Door to Stores  ·  
Home/AI Models/ChatGPT
OpenAI Answers Opus 5.5 With GPT-6 Sol and Luna, at Half the Price

OpenAI Answers Opus 5.5 With GPT-6 Sol and Luna, at Half the Price

ChatGPT

OpenAI launched GPT-6 Sol and GPT-6 Luna 90 minutes after Claude Opus 5.5, cutting API prices roughly in half, though its own benchmarks show the new models scoring lower than their predecessors on some evaluations.

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, roughly 90 minutes after Anthropic shipped Claude Opus 5.5. Sol costs 2 dollars per million input tokens and 10 dollars per million output tokens. Luna costs 0.10 and 0.50. Both are about half the price of the GPT-5.6 models they replace, and OpenAI told VentureBeat the pricing is permanent rather than a launch promotion.

Where do Sol and Luna sit in the lineup?

Beneath GPT-6 Astra, the 10 dollar / 50 dollar flagship released September 3 alongside a company statement that we are now in the AGI era. Until this launch, Astra was the only GPT-6 model, and developers doing everyday work were still on GPT-5.6. Sol is positioned for complex coding and agentic workflows; Luna is built for focused tasks that need to run cheaply at very high volume. OpenAI says both were trained with methods similar to Astra’s.

Both carry a 1.05 million-token context window with input capped at 922K, 128K max output tokens, text and image input, and OpenAI’s agent tooling including function calling, web search, file search, computer use, and MCP connections. Model IDs are gpt-6-sol and gpt-6-luna. One oddity: Luna has a newer knowledge cutoff (May 18, 2026) than either Sol (April 20) or Astra (April 30), unusual for the cheapest model in a family.

Are they actually better, or just cheaper?

Mostly cheaper, and OpenAI is fairly direct about that: its own launch page states that Astra remains its best model across the board. The pitch is cost efficiency, not a new capability ceiling.

The numbers support a mixed reading. On AutomationBench, Sol at xhigh effort scores 33.2 percent at 0.27 dollars per task, beating Claude Opus 5 at max (26.9 percent at 11.1 times the cost) and GPT-6 Astra at low (30.3 percent at 3.9 times the cost). On Agents’ Last Exam, Sol at max scores 56.4 percent against Opus 5’s best of 55.9 percent. But on DeepSWE and OSWorld 2.0, GPT-6 Sol’s best scores (68.8 percent and 64.4 percent) fall below GPT-5.6 Sol’s best (72.7 percent and 66.2 percent) and below Claude Opus 5’s best (73.7 percent and 70.2 percent). On two of six headline evaluations, the new model scores lower than the one it replaces while costing roughly 60 percent less per task.

One comparison deserves scrutiny. OpenAI says Sol can match Claude Fable 5.1 xhigh at much lower cost, and it does: 49.3 percent for 2.14 dollars against 48.7 percent for 9.27 dollars. But xhigh is Fable 5.1’s weakest setting on that chart. Fable 5.1 at low scores 49.8 percent for 2.38 dollars, which is both a higher score and a near-identical price.

What about the accuracy claim?

OpenAI says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol, approaching Astra-level reliability at much lower cost. That figure comes from an internal factuality evaluation using de-identified real-world conversations where users flagged model errors. It is a meaningful methodology, but it is OpenAI’s own test rather than an independent benchmark, and the 50 percent price-cut headline is measured against GPT-5.6 promotional rates, so compare it with what you actually paid.

Why Luna may matter more than Sol

At 0.10 dollars per million input tokens and 0.50 per million output, Luna’s output price fell further than the 50 percent headline suggests, down from 1.20 dollars. That puts classification, extraction, routing, and other high-volume work into territory where inference cost stops being the thing that decides whether a feature can ship profitably. It is the same competitive pressure driving DeepSeek’s aggressive V4.1 Flash pricing and the enterprise shift toward cheaper models documented in our report on AT&T routing 40 percent of its AI traffic to open models.

Availability is narrower than usual at launch. Both models are rolling out in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Free and Go subscribers can reach Luna through the ChatGPT desktop app. Neither is in the regular Chat interface yet.

See VentureBeat’s launch coverage for more.

Up Next
Claude Opus 5.5 Matches Fable 5.1 at 40% Lower Cost

Claude Opus 5.5 Matches Fable 5.1 at 40% Lower Cost

Claude

Anthropic released Claude Opus 5.5, the first model in its 5.5 family, matching Fable 5.1 performance on most work while costing 40 percent less to run and generating output 30 percent faster than Opus 5.

Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new 5.5 family. It performs at the level of Claude Fable 5.1 on most work while costing 40 percent less to run than Opus 5, and generates output more than 30 percent faster. API pricing drops to 4 dollars per million input tokens and 20 dollars per million output tokens, down from 5 and 25. It is available on all platforms, including AWS, Google Cloud, and Microsoft Azure.

What actually improved?

On benchmarks Anthropic published, Opus 5.5 surpassed Fable 5.1 on agentic coding, knowledge work, computer use, visual chart recognition, and multidisciplinary reasoning. The named benchmarks include Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0, and GDPval-AA v2.1. Anthropic cites one early tester completing a 680,000-line code migration in under a day, work it says would have taken an engineering team weeks.

There is also a communication change that is easy to overlook. Anthropic says Opus 5.5 uses less jargon and puts important information at the start of a response rather than building to it, so answers are easier to scan. That is a behavioural change rather than a capability one, but for anyone reading model output all day it may matter more than a benchmark point.

How does the pricing actually break down?

  • Input: 4 dollars per million tokens, down 20 percent from 5 dollars
  • Output: 20 dollars per million tokens, down 20 percent from 25 dollars
  • Cache write: 5 dollars per million tokens, down from 6.25 dollars
  • Cache read: 0.20 dollars per million tokens, down sharply from 0.50 dollars

The headline 40 percent figure is not the same as the 20 percent rate cut. Anthropic arrives at 40 percent by combining the lower rates with the claim that the model uses fewer tokens to finish a given task. That is a real effect if it holds on your workload, but it is a per-task estimate rather than a posted price, so teams should measure it against their own usage rather than assume it.

Why does the safety framing matter here?

Anthropic explicitly positions this as its first release since calling for pacing the frontier, referring to CEO Dario Amodei’s essay urging the industry to slow capability advancement, which we covered in our report on the slowdown call and the safety resignations that preceded it. The company says Opus 5.5 was tested before release by external evaluators including Frontier Design and METR, and that on its automated behavioral audit, the most comprehensive alignment test it runs, Opus 5.5 is the strongest-performing model it has tested.

The obvious tension is that shipping a more capable model two months after Opus 5 is not, on its face, slowing down. Anthropic’s answer is that pacing refers to the rate of capability advancement paired with safety work, not a freeze on releases. Whether that distinction satisfies the researchers who resigned over exactly this question is a separate matter, and one worth watching as Sonnet 5.5 and Haiku 5.5 arrive.

Because Opus 5.5 is comparable to Mythos 5.1 in biology and cybersecurity capability, it ships under the same safeguards applied to Fable 5.1. Those limit how far the model can be pushed on discovering exploits in compiled programs or developing recognizable biological weapons, the same restricted-access logic behind Google’s decision to gate its most capable cybersecurity model.

What changes for subscribers?

Anthropic raised the five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans alongside the launch. Subscribers also received a rate limit reset usable any time between now and October 22. Sonnet 5.5 and Haiku 5.5 are promised in the coming weeks with many of the same improvements to performance, efficiency, and safety.

What is the honest caveat?

Every benchmark figure above comes from Anthropic’s own evaluations. Independent verification has not caught up yet, and the industry has a pattern of vendor-published numbers looking stronger than third-party results, a problem our guide to how AI model rankings work covers in detail. The external evaluator involvement from METR and Frontier Design applies to safety testing, not to the performance claims. Treat the capability numbers as a starting point for your own testing rather than a settled result.

See Anthropic’s own announcement and TechCrunch’s coverage for the full detail.