The Agentic Post
Breaking
Digital Twins and Physical AI  Â·  Humanoid Robots in Manufacturing  Â·  AI Data Centers and Water Usage  Â·  The AI Chip Supply Chain, Explained  Â·  What Is Fine-Tuning? A Plain Explainer  Â·  Meta and Sierra Want to Give AI Agents a Front Door to Stores  ·  
Home/Guides/Comparisons
Claude Opus 5.5 vs GPT-6 Sol: Which Should You Use?

Claude Opus 5.5 vs GPT-6 Sol: Which Should You Use?

Comparisons

Claude Opus 5.5 and GPT-6 Sol launched 90 minutes apart. A practical comparison of pricing, benchmark results, context windows, and which workloads each model actually fits.

Claude Opus 5.5 and GPT-6 Sol launched within 90 minutes of each other on September 22, 2026. Opus 5.5 costs 4 dollars per million input tokens and 20 per million output. GPT-6 Sol costs 2 and 10, exactly half. Neither is straightforwardly better: Opus 5.5 aims at the top of the capability range, Sol aims at the best result per dollar. Which one fits depends almost entirely on whether your bottleneck is capability or cost.

How do the prices compare?

  • Claude Opus 5.5: 4 dollars input / 20 dollars output per million tokens. Cache write 5 dollars, cache read 0.20 dollars.
  • GPT-6 Sol: 2 dollars input / 10 dollars output per million tokens. Cached input 0.20 dollars.
  • GPT-6 Luna: 0.10 dollars input / 0.50 dollars output, for high-volume work that does not need either model’s reasoning.

Raw rates understate the picture on both sides. Anthropic claims typical workloads cost about 40 percent less than Opus 5 overall, because Opus 5.5 also uses fewer tokens to finish a task, so the effective gap is narrower than 2x. OpenAI’s cost-per-task charts similarly show Sol finishing work far cheaper than its own predecessor. Per-task cost, not per-token price, is the number that actually decides your bill.

Which one performs better?

On the vendors’ own numbers, Opus 5.5 sits higher. Anthropic reports it surpassing Claude Fable 5.1 on agentic coding, knowledge work, computer use, visual chart recognition, and multidisciplinary reasoning, and calls it the strongest model it has tested. OpenAI, by contrast, states plainly on its own launch page that GPT-6 Astra remains its best model across the board, positioning Sol as a cost-efficiency release rather than a capability ceiling.

The detail worth knowing: on two of six headline evaluations, GPT-6 Sol’s best score is actually lower than GPT-5.6 Sol’s best. On DeepSWE it scores 68.8 percent against 72.7 percent, and on OSWorld 2.0 64.4 percent against 66.2 percent. It is cheaper per task, not uniformly stronger. Where Sol does win is cost-adjusted: on AutomationBench it scores 33.2 percent at 0.27 dollars per task, beating Claude Opus 5 at max effort (26.9 percent) at roughly one eleventh the cost.

What about context and speed?

GPT-6 Sol carries a 1.05 million-token context window with input capped at 922K and up to 128K output tokens. Anthropic reports Opus 5.5 generating output more than 30 percent faster than Opus 5. Both support the standard agent tooling: function calling, web search, file search, computer use, and MCP connections on OpenAI’s side.

So which should you actually pick?

Choose Opus 5.5 if your work is long-running coding agents, complex refactors, or knowledge tasks where a wrong answer costs more than the tokens do. The 680,000-line code migration Anthropic cites as an early tester result is the shape of problem it is built for.

Choose GPT-6 Sol if you are running high volumes of moderately hard work and cost per task is what determines whether the product is viable. Choose GPT-6 Luna for classification, extraction, and routing, where 0.10 dollars per million input tokens changes what is economically possible at all.

Consider neither if your workload is mostly routine. The enterprise pattern documented in our report on AT&T routing 40 percent of its AI traffic to open models found a 56 percent cost cut for a 2 percent quality drop, and DeepSeek V4.1 Flash undercuts both on cached input. Routing by task rather than standardising on one model is usually the cheaper answer.

The caveat that applies to both

Every number above comes from the vendor that sells the model. Independent benchmarking has not caught up with either release, and as our guide to reading AI model rankings explains, vendor-selected comparisons routinely flatter the vendor. One example from this launch: OpenAI’s claim that Sol matches Claude Fable 5.1 at much lower cost uses Fable’s weakest setting on that chart. Run both against your real workload for a week before committing. Anthropic’s Opus 5.5 announcement and VentureBeat’s GPT-6 coverage hold the source figures. For the full picture see our coverage of the Opus 5.5 launch and the GPT-6 Sol and Luna release.

Up Next
Ex-Anthropic Startup Mirendil Seeks $5B Valuation in Three Months

Ex-Anthropic Startup Mirendil Seeks $5B Valuation in Three Months

Funding & Startups

Mirendil, founded by former Anthropic researchers to build AI that automates AI research, is in talks to raise up to $1 billion at a $5 billion valuation, five times its seed valuation from three months ago.

Mirendil, the AI startup founded by former Anthropic researchers, is in talks to raise up to 1 billion dollars at a roughly 5 billion dollar valuation, according to Bloomberg. Kleiner Perkins is in discussions to lead, with Andreessen Horowitz also in talks to participate. The round comes three months after the same two firms led a 200 million dollar seed at a 1 billion dollar valuation, a fivefold valuation jump in a single quarter.

Who is behind Mirendil?

Behnam Neyshabur (CEO) and Harsh Mehta (CTO), both former Anthropic researchers, founded the company in December 2025. Neyshabur co-led Anthropic’s scientific AI reasoning team and previously spent more than five years at Google DeepMind, where he co-led reasoning research for Gemini. Mehta was a senior research scientist at Anthropic. The two first met at Google in 2019. The founding team also includes Shayan Salehian, an early member of xAI, and Tara Rezaei, an MIT graduate and former OpenAI intern. The team now numbers around 20 researchers drawn from Anthropic, xAI, DeepMind, and OpenAI.

What is the company actually building?

AI that does the work of an AI researcher. Mirendil trains specialised models on frontier AI research tasks, including experimental design, hyperparameter search, model evaluation, and iterative training, then packages those capabilities as a platform other organisations can deploy on their own problems. The stated goal is democratising frontier research: a university biology lab could use it to build and refine a model for drug-target identification without a dedicated machine learning engineering team.

To support the compute demands of training self-improving models, Mirendil signed a 100 million dollar Google Cloud deal in August 2026, giving it access to TPUs, NVIDIA GPUs, and managed clusters. If the current round closes, the company will have raised roughly 1.2 billion dollars within about nine months of founding.

Why is this timing uncomfortable?

Because AI that improves AI is precisely the capability Anthropic’s CEO named as a reason to slow down. In the essay we covered in our report on the industry slowdown call, Dario Amodei cited the increasing ability of AI systems to build more advanced AI as one of two developments that changed his thinking, because it makes capability gains compound rather than accumulate linearly. Mirendil is a company built to accelerate exactly that loop, funded by investors who also back the labs, and staffed substantially by people who left those labs.

That is not a criticism of Mirendil’s stated mission, which is about widening access to research capability rather than racing ahead of it. But the two things sit in genuine tension, and the money is currently flowing toward acceleration at a speed that dwarfs anything happening on the governance side.

What does the valuation actually signal?

Mirendil has no publicly disclosed revenue and a product still being built. A 5x valuation increase in three months reflects investor conviction about the category and the founding team’s pedigree, not demonstrated commercial traction, a dynamic our explainer on how AI startup valuations are set covers in more depth. It also fits the broader neo-lab pattern: experienced researchers leaving major labs to found focused startups that raise enormous sums almost immediately.

Terms have not been finalised and the round has not closed. Both figures come from people familiar with the discussions speaking anonymously, so treat them as reported rather than confirmed.

See Bloomberg’s original report.