The Agentic Post
Breaking
Gemini’s Multimodal Features, Explained  Â·  ChatGPT Custom GPTs, Explained  Â·  What Is Constitutional AI? Explained  Â·  AI Capex Explained for Investors  Â·  AI Startup Valuations: How They Are Set  Â·  How to Reskill for an AI Job Market  ·  
Home/AI Models/ChatGPT
ChatGPT Goes Unlimited for Free Users

ChatGPT Goes Unlimited for Free Users

ChatGPT

OpenAI removed ChatGPT's free-tier text chat limit and updated its default models, unifying instant and reasoning modes for paid users while giving free users unlimited access to GPT-5.6 Luna.

OpenAI removed the text chat limit on ChatGPT’s free tier this week, alongside a broader update that reworks which model powers the product for every paying and non-paying user. The changes land five weeks after OpenAI’s GPT-5.6 family reached general availability, and they arrive with a specific number attached to them: ChatGPT recently crossed one billion weekly users, a milestone OpenAI referenced directly when explaining the update.

What actually changed for each tier

Free and Go plan users are moving to GPT-5.6 Luna as their new default model, replacing GPT-5.5 Instant, and gaining unlimited text conversations without the rate limit that previously forced a wait once a session cap was reached. They are also getting a new Think button that triggers deeper reasoning for harder questions, subject to anti-abuse guardrails OpenAI did not detail publicly. Limits remain in place for everything outside plain text chat: file uploads, image generation, and voice all keep their existing caps.

Plus and Pro subscribers get a different change. Rather than switching between a separate Instant model and a separate reasoning model depending on the task, both are now handled by a single, updated GPT-5.6 Sol model, with a new slider letting users manually control how much computational effort ChatGPT puts into each response, from fast everyday answers up through more thorough analysis for coding, planning, and research. OpenAI says the goal was a model that delivers focused answers, adapts its level of detail to the question, and avoids unnecessary formatting regardless of how much reasoning effort it applies, eliminating what the company described internally as a jarring shift in tone that previously occurred when a conversation crossed from quick-answer mode into thinking mode. Notably, this update is scoped narrowly: the version of GPT-5.6 Sol used inside ChatGPT Work and Codex is explicitly unchanged, only the consumer chat experience is affected.

The accuracy numbers OpenAI is citing

OpenAI says internal testing across financial, medical, and legal prompts found factual errors dropped by roughly 62 percent with the new GPT-5.6 Luna and 68 percent with the updated GPT-5.6 Sol, compared against the prior GPT-5.5 Instant model. Those are OpenAI’s own internal figures rather than independently verified benchmark results, and the company has not published the specific test set or methodology behind them, so they’re worth treating as a company claim rather than a settled, externally reproduced number. Still, directionally they target exactly the failure mode that matters most for a model serving over a billion people weekly: confidently wrong answers on questions involving specific dates, numbers, sources, or rules, precisely the territory where AI hallucination tends to concentrate and where getting it wrong carries real consequences.

Why unlimited free access is a retention play, not a cost play

Removing the text-chat cap for free users is worth reading as a top-of-funnel and retention move rather than a cost-cutting one. It arrives after a year in which OpenAI has actually tightened usage limits in other parts of the product, and running an always-available free tier at over a billion weekly users is not a cheap decision from a compute-cost standpoint. The logic instead appears to be about keeping the free tier genuinely useful enough that people stay inside the ChatGPT ecosystem rather than bouncing to a competitor once they hit a wall mid-conversation, a friction point that has reportedly driven real churn toward rivals offering more generous free access.

The rollout itself is staggered. GPT-5.6 Luna became the default model for Free and Go users starting the week of the announcement, while the unlimited text-chat change and the new Think button followed roughly a week later. For Plus and Pro users, the updated GPT-5.6 Sol model and the new reasoning-effort slider began rolling out immediately across ChatGPT’s web, mobile, and desktop apps.

Part of a broader Sol, Terra, Luna naming shift

This update is a consumer-facing simplification pass rather than a new model generation. The underlying Sol, Terra, and Luna tier system was introduced in June, with Sol reaching general availability in July, and this week’s changes fold what had been a confusing choice, between separate Instant and Thinking modes, into a single model with adjustable effort instead. It also lands a month after OpenAI cut prices on the Luna and Terra tiers in late July, meaning the free and mid-tier models have now seen both a price cut and a capability upgrade within a matter of weeks, a combination that suggests OpenAI is prioritizing broad accessibility across its full user base rather than concentrating improvements only at the top of its subscription ladder.

How this compares to what rivals offer

The competitive backdrop matters here. Google’s Gemini already offers a genuinely capable free tier woven directly into Search, Gmail, and Docs, and Anthropic and xAI have both leaned on generous free access as part of their own growth strategies this year. A free tier that cuts a user off mid-conversation is a real point of friction in a market where switching to a competitor costs nothing but a new tab. Removing that friction for OpenAI’s largest, least monetized user segment is a defensive move as much as an offensive one: it protects the base ChatGPT has built rather than only trying to convert more of it to paid tiers.

The real cost this shifts, not eliminates

None of this makes free-tier inference free to run. OpenAI is absorbing real, ongoing compute cost to serve unlimited text conversations at a user base measured in the billions, a bet that only pays off if enough of those free users either convert to paid tiers over time or generate enough platform value, through engagement, data, and ecosystem lock-in, to justify the spend on their own. That calculation sits inside a much larger one: OpenAI’s overall infrastructure spending has scaled dramatically alongside its user growth this year, and a free tier this generous is only sustainable if the company’s broader unit economics continue trending the direction OpenAI’s public statements suggest they are.

What to actually watch next

The clearest signal of whether this update succeeds on its own terms will show up in two places over the coming weeks: whether OpenAI’s own accuracy claims hold up under independent testing once enough people have used the updated models at scale, and whether daily active usage among free-tier users actually increases now that the session-ending wall is gone. If the unlimited-access change meaningfully increases how often free users return, that’s a strong signal the friction really was costing OpenAI retention. If usage patterns stay roughly flat, it suggests the previous limits weren’t the binding constraint on engagement that this update assumes they were, and OpenAI absorbed a real, ongoing compute cost for a change that mostly benefits people who were already using the product heavily within the old limits anyway.

See Help Net Security’s original coverage for additional detail on the safety guardrails accompanying the rollout.

Up Next
Anthropic Confirms Its Own AI Chip Team

Anthropic Confirms Its Own AI Chip Team

Chips & GPUs

Anthropic confirmed it is building an in-house team to co-design custom silicon for Claude, aiming to cut inference costs while maintaining its existing multi-chip strategy across AWS, Google, Nvidia, and AMD.

Anthropic confirmed this week that it is building an in-house team to design custom silicon for Claude, joining a small but growing group of AI labs betting that owning part of the hardware stack is now worth the enormous cost of doing it. The company is hiring engineers spanning hardware and software to co-design chips and models together, with postings describing the group as a custom silicon team and offering salaries reportedly ranging from 320,000 to 485,000 dollars for candidates who have personally shipped finished semiconductor designs.

What Anthropic actually said, and did not say

The confirmation itself was narrow. Anthropic said it is assembling engineers who will develop chips and AI models in tandem, aiming to run Claude faster and more cost-efficiently at the scale its customers now require. It offered no timeline for when the effort might produce a finished part, no confirmed manufacturing partner, and no indication of whether it intends to handle fabrication itself at all. Reports earlier in the summer suggested Anthropic had been in talks with Samsung as a potential manufacturing partner, but that remains unconfirmed by the company. Every part of the announcement that would represent a genuine break from Anthropic’s existing supply chain is, notably, still missing.

Technical leadership of the new effort has reportedly landed with Clive Chan, who joined Anthropic in early June. Chan was previously the second hardware hire on OpenAI’s own dedicated chip team, which he joined from Tesla’s Dojo supercomputer program, where he worked on GPU optimization and training infrastructure for Autopilot. At OpenAI he worked on the matrix multiplication architecture behind that company’s own in-progress chip, reportedly already at prototype stage, before moving to Anthropic’s newer, earlier-stage program. In a post announcing the move, Chan described the shift as climbing a new mountain from the base, a fair summary of the gap between a chip that already exists in prototype form and one that has only just been formally named.

Why the economics finally tipped this way

Designing an advanced AI chip from scratch is not a decision labs make lightly. Industry estimates cited by Reuters put the cost of developing a single advanced AI chip at close to half a billion dollars, driven by specialized engineering talent and the difficulty of achieving reliable, defect-free fabrication at leading-edge nodes. That number only makes sense to spend if the resulting savings, multiplied across enough usage, clear the bar. Anthropic appears to believe it now does. The company is reportedly targeting roughly a 50 percent reduction in per-token inference cost through co-designing silicon and models together, tailoring chip architecture directly to the attention mechanisms Claude actually uses rather than running on general-purpose hardware built for a broader range of workloads.

The backdrop that makes this economically rational is a genuinely dramatic cost curve. The per-query cost of running a model roughly equivalent to GPT-3.5 fell from about 20 dollars per million tokens in late 2022 to around 7 cents per million tokens within a few years, according to figures Anthropic has cited. At that scale of decline, and at the volume of queries a frontier lab now serves, even a chip that costs hundreds of millions of dollars to design can pay for itself many times over if it meaningfully improves cost per token at Anthropic’s actual usage levels. Custom AI accelerator shipments are on track to grow roughly 44.6 percent in 2026, according to TrendForce data, nearly three times the 16.1 percent growth rate projected for general-purpose GPUs, a trend line that reflects exactly this logic playing out across the industry: for workloads that are predictable and stable at high enough volume, a custom chip’s design cost is straightforwardly amortized.

Multi-chip, not single-chip

Anthropic has been explicit that this does not replace its existing hardware relationships. The company said it plans to maintain a multi-chip strategy, continuing to run on hardware from AWS, Google, Nvidia, and AMD alongside anything its own team eventually produces. That framing matters: Anthropic did not wait for an in-house chip team to start diversifying its compute. In April, the company expanded its partnership with Google and Broadcom for roughly 3.5 gigawatts of next-generation TPU capacity coming online in 2027, on top of a gigawatt already arriving in 2026 under a Google Cloud agreement signed the previous October. Anthropic is also at the center of a 15 billion dollar financing deal for an AI data center campus in Hubbard, Texas, being developed by Nexus Data Centers, where Google is backstopping Anthropic’s obligations and the company intends to deploy TPUs co-developed by Google and Broadcom under a vendor-financing arrangement.

Seen against that backdrop, the custom silicon team looks less like a bid to escape dependence on any single supplier and more like Anthropic adding one more lever to an already diversified hardware strategy, one it now controls end to end rather than negotiating through a partner. Google itself is not a neutral bystander here either: it has agreed to invest up to 40 billion dollars in Anthropic, giving it a direct financial stake in Anthropic’s success even as Anthropic works to reduce its dependence on any one hardware vendor, Google included.

The physics Anthropic still cannot design around

However this program develops, it runs into the same underlying constraint every AI hardware effort eventually meets: a handful of foundries control leading-edge manufacturing capacity worldwide, and the lithography equipment required to produce advanced chips can cost up to 400 million dollars per machine. Designing your own chip frees a company from depending on a single chip designer. It does not free it from depending on a foundry, and it does not free it from the physical limits of current manufacturing capacity, both of which remain genuine bottlenecks regardless of who owns the design.

Anthropic is not alone in reaching this conclusion. AMD and Meta signed a multi-year chip supply agreement reportedly worth as much as 60 billion dollars earlier this year, and OpenAI’s own chip program, the one Chan left to join Anthropic, is reportedly already further along, with a prototype in hand. The pattern across frontier labs is now consistent: as usage scales into the volumes these companies are now serving, the economics of custom silicon increasingly favor building rather than only renting, even for companies that started, and remain, entirely dependent on someone else’s fabs to actually manufacture anything.

Why co-design is the actual bet, not just cost

The distinction Anthropic keeps emphasizing, that engineers will develop chips and models in tandem rather than sequentially, is worth taking seriously as the real strategic logic here, separate from the headline cost savings. A general-purpose GPU has to serve an enormous range of workloads well, which means real efficiency is left on the table for any single, specific use case. A chip designed alongside the model that will actually run on it can be shaped around the exact computational patterns Claude’s architecture produces, rather than the broader range of patterns a general-purpose accelerator has to handle adequately. That’s a genuinely different design philosophy than simply buying more of an existing chip, and it’s the same logic that has made Google’s TPUs, developed over more than a decade specifically for its own workloads, a real efficiency advantage inside Google’s own infrastructure.

The talent war this signals

The salary range attached to Anthropic’s job postings, up to 485,000 dollars for engineers who have personally shipped finished chip designs, is itself a signal worth reading. Semiconductor design talent capable of taking a chip from concept to working silicon is genuinely scarce, and every frontier AI lab now competing for the same small pool of engineers is bidding those numbers up in the process. Chan’s own path, from Tesla’s Dojo program to OpenAI’s chip team to Anthropic in the space of roughly two years, illustrates how mobile that talent pool has become, and how directly labs are now recruiting engineers away from each other’s hardware programs specifically, not just their model research teams.

See Forbes’ original reporting for more on how this compares to rival labs’ chip programs.

This connects directly to the same underlying manufacturing economics covered in our explainer on how AI chips are actually made.