A viral post claiming Alibaba "just launched" Qwen3-Max at 6 dollars per million tokens against Claude at 25 has been circulating this week. Almost every number in it is wrong or a year out of date. But the underlying question is real, because Alibaba’s actual current flagship, Qwen3.8-Max, costs 2 dollars input and 6 dollars output per million tokens against Claude Opus 5.5 at 4 and 20. Here is the accurate comparison.
First, what the viral post got wrong
- Qwen3-Max did not just launch. It shipped September 23, 2025, more than a year ago. Alibaba is three generations past it: Qwen3.7-Max in May 2026, Qwen3.8-Max in August 2026.
- The 6 dollar figure is not Qwen3-Max’s price. Qwen3-Max runs about 0.78 input and 3.90 output. The 6 dollar output figure belongs to the newer Qwen3.8-Max.
- The 25 dollar Claude figure is stale. That was Opus 5. Opus 5.5 launched September 22 at 20 dollars output.
- "Rivals Opus 5 on complex reasoning" overstates it. Qwen3-Max scores 19.4 on the Artificial Analysis intelligence index, around the 56th percentile. Not frontier.
What do these models actually cost?
- Qwen3.8-Max: 2 dollars input, 6 output. Implicit cache reads 0.25, explicit cache reads 0.17
- GPT-6 Sol: 2 input, 10 output, cached input 0.20
- Claude Opus 5.5: 4 input, 20 output, cache reads 0.20
- Claude Sonnet 5.5: 2 input, 10 output, cache reads 0.20
So the real gap between Alibaba’s flagship and Anthropic’s is roughly 3.3x on output, not the 4.17x the viral graphic claimed. And Sonnet 5.5 matches Qwen on input price exactly.
Does Qwen3.8-Max actually match the frontier?
On some things yes, on the hardest things no. Alibaba’s own published table shows the split clearly. On Terminal-Bench 2.1, Qwen3.8-Max scores 86.6 against 84.6 for both Opus 4.8 and Fable 5, and 88.8 for GPT-5.6 Sol. That is a genuine win over two Claude flagships.
On SWE-bench Pro it scores 67.7 against Fable 5’s 80.0, a 12-point gap on core software engineering. On DeepSWE 1.1 it posts 56.6 against Fable 5’s 70.0 and GPT-5.6 Sol’s 73.0. The pattern holds across the table: competitive on terminal-driven agentic tasks, behind on deep engineering.
It also wins outright in places worth knowing about: 82.8 on IFBench instruction following against 72.7 for GPT-5.6 Sol, 73.2 on PLawBench, 60.2 on HealthBench, and 86.1 on OSworld-Verified for desktop agent work. GPQA Diamond is effectively a tie at 92.6.
The benchmark caveat that matters most
Alibaba ran these evaluations itself, on its own infrastructure, and reports the highest score among harnesses. It notes that Qwen3.8-Max performs best on Claude Code, which is the harness used for its headline Terminal-Bench number. Comparison figures for rivals come from different harnesses and different runs. This is normal practice and it is also why vendor tables consistently flatter the vendor, a problem our guide to model rankings covers.
Note also that the comparison set is Opus 4.8 and Fable 5, not Opus 5.5. On the newer Terminal-Bench 4.0, Opus 5.5 leads at 66.4% while Qwen3.8-Max sits around 67.7 on a differently-run board. Cross-version comparisons here are not clean.
The thing the viral post missed entirely
Alibaba released the weights. Qwen3.8-2.4T-A95B went public on August 12, 2026: 2.4 trillion parameters with 95 billion active per token, 262K native context extensible to about 1.01 million. The open build omits image input and non-thinking mode, but it is a genuinely frontier-adjacent model you can self-host.
That matters more than a price comparison, because it removes the vendor from the equation entirely for regulated environments that cannot call third-party APIs. It is the same dynamic behind AT&T routing 40% of its AI traffic to open models and DeepSeek shipping V4.1 Flash under MIT.
So which should you use?
Qwen3.8-Max for high-volume agentic and terminal work, desktop automation, instruction-heavy pipelines, or anywhere self-hosting is a hard requirement. Opus 5.5 for deep engineering, long refactors, and work where a wrong answer costs more than the tokens. GPT-6 Sol when cost per task is the deciding constraint. Sonnet 5.5 is the one the viral post ignored and arguably the most direct competitor on price.
The honest summary: the price gap is real and narrowing the quality gap is real too, but a 3.3x price difference is not the same as equivalence, and a year-old model quoted at the wrong price is not evidence of either. See Alibaba’s own benchmark table and the open weights on Hugging Face.




