The Agentic Post
Breaking
Gemini’s Multimodal Features, Explained  Â·  ChatGPT Custom GPTs, Explained  Â·  What Is Constitutional AI? Explained  Â·  AI Capex Explained for Investors  Â·  AI Startup Valuations: How They Are Set  Â·  How to Reskill for an AI Job Market  ·  
Home/AI Tools/Coding Assistants
DeepSeek’s New Model Costs Pennies

DeepSeek’s New Model Costs Pennies

Coding Assistants

DeepSeek-V4-Flash-0731 exited preview at $0.14/$0.28 per million tokens, beating DeepSeek's own 1.6-trillion-parameter Pro model on agentic coding benchmarks.

DeepSeek pushed AI coding costs closer to zero on August 1, 2026, releasing DeepSeek-V4-Flash-0731 out of preview at $0.14 per million input tokens and $0.28 per million output tokens, pricing that undercuts most Western competitors by an order of magnitude while beating DeepSeek’s own larger flagship model on agentic coding benchmarks.

Quick facts

  • DeepSeek-V4-Flash-0731 exited preview on August 1, 2026, priced at $0.14 per million input tokens and $0.28 per million output tokens.
  • It scores 82.7% on Terminal-Bench 2.1 and 54.4% on DeepSWE, beating DeepSeek’s own 1.6-trillion-parameter V4-Pro-Preview model on agentic coding benchmarks.
  • It ships with native Responses API support and Codex compatibility, plus DeepSeek’s own speculative decoding stack for 2-3x token throughput on H800 hardware.
  • Legacy deepseek-chat and deepseek-reasoner API aliases have been retired in favor of deepseek-v4-flash.
  • DeepSeek-V4-Pro’s full general availability remains pending, reportedly delayed into the August 10-20 window.

Why a smaller model beating a bigger one is the actual story

Per Axios’s reporting, the significant part isn’t just the price, it’s that a lighter, cheaper “Flash” model outperforming DeepSeek’s own much larger Pro-tier model on real coding benchmarks confirms something the industry has been circling for a while: post-training technique and data quality are starting to matter more than raw parameter count for practical coding tasks. That’s a direct threat to the pricing power of every lab still charging a premium primarily on the basis of model size.

The race to zero is now industry-wide, not just DeepSeek

DeepSeek isn’t cutting prices in isolation. OpenAI cut its own budget-tier Luna model from $1/$6 to $0.20/$1.20 per million tokens on July 30, 2026, and Anthropic’s Claude Sonnet 5 is running introductory pricing specifically to stay competitive during this window. Every major lab is now defending against the same pressure: a genuinely capable model at a fraction of the price forces competitors to either match on price or clearly justify a premium with capability that Flash-tier pricing can’t touch yet, like the hardest reasoning, judgment, and cybersecurity tasks.

What this actually means for anyone building on it

For agent and application builders running high-volume coding tasks, deepseek-v4-flash’s combination of price and Terminal-Bench score makes it a genuinely serious default option, not just a cheap fallback, for workloads where DeepSeek’s benchmarks hold up against your actual codebase. The retirement of the legacy deepseek-chat and deepseek-reasoner aliases means anyone still pointed at those needs to migrate their API calls to deepseek-v4-flash directly, a small technical task worth doing sooner rather than later given the old aliases are no longer the current model.

Common questions

Is DeepSeek-V4-Flash open-weight? DeepSeek’s prior model generations have been released as open-weight; the reporting reviewed here covers the API release specifically and doesn’t confirm an open-weight release date for this exact checkpoint.

Do I need to change my code to use it? If you’re using the legacy deepseek-chat or deepseek-reasoner model aliases, yes, those are retired and calls should point to deepseek-v4-flash directly. New integrations should use the current model name from the start.

When is DeepSeek-V4-Pro coming out? Reporting points to a general availability window between August 10 and August 20, 2026, though DeepSeek has not confirmed an official date; treat any specific date as unofficial until DeepSeek’s own channels confirm it.

Key takeaway

If cost per token is a meaningful line item in your AI spend, DeepSeek-V4-Flash-0731 is worth benchmarking against your actual workload now, not waiting for V4-Pro’s delayed general release. The gap between this model’s price and its coding benchmark scores is large enough that it changes the math for high-volume use cases specifically, even if it doesn’t touch the hardest reasoning tasks frontier labs still charge a premium for.

Up Next
Big Tech AI Spending Splits Winners

Big Tech AI Spending Splits Winners

Big Tech

Microsoft, Amazon, and Alphabet gained nearly $1.5 trillion in market value during earnings week while Apple, Meta, and Tesla lost ground, as investors demand AI spending show up as revenue.

Big Tech’s earnings week just rewrote the market’s story on AI spending. Amazon, Microsoft, and Alphabet added nearly $1.5 trillion in combined market value in a single week, while Apple, Meta, and Tesla lost a combined chunk of that swing, over $2 trillion moved in or out of six companies in a matter of days.

Quick facts

  • Microsoft gained more than $600 billion in market value during earnings week; Amazon and Alphabet each added more than $400 billion, per CNBC data.
  • Apple shed more than $350 billion in market cap after supply problems dimmed its outlook.
  • Meta lost more than $85 billion as investors questioned returns on its AI spending; Tesla dropped over $7 billion after posting negative free cash flow.
  • Combined Big Tech AI capital spending is on track to reach roughly $800 billion over the next 12 months, per Jefferies estimates.
  • Amazon raised its 2026 capex forecast to $220 billion and still saw its stock rise.

Why the market is splitting winners from losers so sharply

Per Yahoo Finance’s reporting, the dividing line isn’t who’s spending the most on AI, it’s who can point to that spending already converting into recurring revenue. Microsoft’s gain was driven by Azure and cloud growth investors could see directly in the numbers; Alphabet’s Google Cloud rounded out what the piece calls a “victorious trifecta” alongside Amazon and Microsoft. All three share something Apple, Meta, and Tesla don’t have at comparable scale: established cloud platforms that can turn chip and data center spending directly into subscription and enterprise AI revenue, rather than absorbing the cost purely as a bet on future products.

The Alphabet whiplash is the clearest example of the new mood

Just a week earlier, the picture looked very different. Fortune reported that Alphabet shares plunged more than 7% in a single day, their worst in over a year, after the company raised 2026 capex guidance to as much as $205 billion and reported negative free cash flow for the first time since its 2004 IPO, despite delivering an 82% jump in cloud revenue that beat Wall Street estimates comfortably. That a company posting genuinely strong cloud growth could still get punished that hard on the same report is exactly the shift in mood driving this whole story: investors have stopped rewarding AI spending on faith alone and started demanding the revenue show up in the same quarter as the capex.

This is the same tension playing out across the chip supply chain

This pattern isn’t isolated to the six companies above. It’s the same dynamic behind the $1 trillion chip stock selloff in late July, where TSMC’s stronger-than-expected earnings still triggered a selloff because the accompanying capex guidance raised fears about margin compression. And it echoes Oracle’s debt-fueled AI infrastructure bet, where the market is specifically scrutinizing whether a company’s spending is backed by cash flow or by leverage. Across chipmakers, cloud giants, and infrastructure-heavy bets alike, the market is applying the same test right now: show the revenue, or get punished for the spending.

Common questions

Does this mean AI spending is slowing down? No. Combined capex is still rising, Amazon and Meta both raised their own 2026 forecasts. What’s changed is how forgiving investors are about spending that isn’t yet showing up as revenue.

Why did Alphabet get punished despite strong cloud growth? The combination of raised capex guidance and negative free cash flow overshadowed the revenue beat for investors focused on near-term cash generation, even though the underlying cloud business grew 82% year-over-year.

What should Meta and Apple do differently? The reporting doesn’t prescribe a fix, but the pattern suggests investors want to see AI investment converting into a clear, reportable revenue line, the way Azure and Google Cloud do, rather than being described mainly in terms of future potential.

Key takeaway

The $800 billion in projected Big Tech AI capex over the next year isn’t going away, but which companies get rewarded for it has clearly changed. Cloud platforms that can point directly to AI-driven revenue growth in the same earnings report are being rewarded; companies asking investors to trust that the spending will pay off later are getting punished immediately, regardless of how strong their underlying AI products actually are.