The Agentic Post
Breaking
Gemini’s Multimodal Features, Explained  Â·  ChatGPT Custom GPTs, Explained  Â·  What Is Constitutional AI? Explained  Â·  AI Capex Explained for Investors  Â·  AI Startup Valuations: How They Are Set  Â·  How to Reskill for an AI Job Market  ·  
Home/Business/Enterprise Adoption
AT&T Sends 40% of Its AI Traffic to Open Models

AT&T Sends 40% of Its AI Traffic to Open Models

Enterprise Adoption

AT&T now routes 40 percent of employee AI requests to open-weight models instead of proprietary AI, cutting coding costs 56 percent for a 2 percent quality drop, illustrating the growing 'good enough' threat open-source AI poses to frontier labs.

AT&T now routes 40 percent of its employees’ AI requests to open-weight models instead of proprietary offerings from OpenAI or Anthropic, according to AT&T vice president Mark Austin, with an internal target of pushing that share to 60 or 70 percent within a few years. The company processes roughly 45 billion tokens a day across its internal AI systems, and Austin says open-source routing has already cut coding-related AI costs by 56 percent, for what the company measured as only a 2 percent drop in output quality. It’s a concrete, quantified example of a shift the New York Times recently framed in stark terms: corporate America is getting hooked on open-source AI, and that shift poses a genuine “good enough” threat to the commercial position of frontier labs like Anthropic and OpenAI.

Why “good enough” is the phrase that matters

The threat frontier labs face here isn’t that open-weight models have caught up to the absolute ceiling of what Claude Opus 5 or GPT-5.6 Sol can do on the hardest, most demanding tasks. It’s that for a large and rapidly growing share of everyday enterprise workloads, routine coding tasks, internal documentation, first-pass customer support drafts, an open-weight model running at a fraction of the cost delivers output indistinguishable enough from a frontier model’s output that the price difference stops being justifiable. AT&T’s own numbers make that calculation concrete: a 56 percent cost reduction against only a 2 percent quality decline is, for the overwhelming majority of business use cases, a straightforwardly rational trade to make, and one that compounds dramatically at AT&T’s actual usage volume of 45 billion tokens daily.

This dynamic mirrors a pattern our own explainer on small versus large AI model tradeoffs covers directly: efficient, purpose-tuned models increasingly close the practical gap with larger, more expensive ones for well-defined, high-volume tasks, even when they remain behind on the hardest, most open-ended problems. What’s new here isn’t the underlying technical trend, it’s a Fortune 100 company putting a specific, verifiable percentage on how far that substitution has already gone inside a single organization’s real production workload.

Why this specifically threatens Anthropic and OpenAI’s business model

Anthropic and OpenAI have both built substantial parts of their revenue on enterprise customers paying premium API rates for access to frontier-tier models. Anthropic in particular has leaned into an enterprise-first strategy, with Menlo Ventures reporting the company holds 42 percent market share in coding applications compared to OpenAI’s 21 percent, and 32 percent of broader enterprise AI usage against OpenAI’s 25 percent. That enterprise concentration is exactly the customer base an internal AT&T-style shift toward open-weight routing puts pressure on directly: the more that large enterprise customers build internal infrastructure to intelligently route routine requests to cheaper open models while reserving frontier models only for tasks that genuinely require their extra capability, the smaller the addressable spend frontier labs can count on from their highest-value customer segment.

That pressure compounds with the broader trend our coverage of DeepSeek’s training efficiency and Meta’s Muse Glimmer release have both tracked this year: the quality gap between leading open-weight models and closed frontier models has been narrowing steadily, even as open models’ cost advantage has stayed roughly constant or widened. AT&T’s routing strategy isn’t a one-off cost-cutting experiment, it’s the kind of infrastructure investment that becomes more valuable, not less, as that quality gap keeps closing.

Why frontier labs aren’t standing still

Both Anthropic and OpenAI have visible countermeasures already in motion. Anthropic’s own revenue growth, reportedly exceeding a 30 billion dollar annualized run rate with year-over-year growth above 1,400 percent according to recent reporting, suggests that even with pressure from open-weight substitution at the margins, aggregate demand for frontier-tier capability is still climbing fast enough to support that growth, at least for now. Both major labs have also leaned harder into tiered pricing and cheaper, faster model variants, exactly the kind of response that acknowledges open-weight competition without conceding the premium end of the market. OpenAI’s recent Ultrafast tier, offering dramatically faster inference at a premium price, is one example of frontier labs trying to justify continued price premiums through genuinely differentiated capability, speed in that specific case, rather than competing purely on raw output quality where open models are closing the gap fastest.

What this means for how enterprises should actually think about AI spend

AT&T’s approach offers a genuinely useful template for other large organizations rethinking their own AI cost structure: rather than treating “which AI vendor should we use” as a single, company-wide decision, route different task categories to different models based on actual measured quality tolerance for that specific task, and revisit that routing regularly as both open and closed model quality keeps shifting. A 2 percent quality drop might be entirely acceptable for internal documentation or routine code review, and completely unacceptable for a customer-facing legal or financial output where an error carries real consequences. The value in AT&T’s disclosed numbers isn’t a blanket argument that open models are now equivalent to frontier ones, it’s a concrete demonstration that intelligent, task-specific routing between the two, rather than an all-or-nothing vendor choice, is where real, substantial cost savings are actually available right now.

Whether frontier labs successfully defend their premium positioning against this kind of substitution, or whether the AT&T pattern becomes the norm across large enterprises over the next several years, is likely to be one of the more consequential open questions shaping AI industry economics through the rest of this decade, well beyond the specific companies and models involved in this particular story.

See AI Weekly’s coverage, drawing on the original New York Times reporting, for additional detail.

Up Next
Google Built an AI Too Dangerous to Release Openly

Google Built an AI Too Dangerous to Release Openly

AI Safety

Google launched Gemini 3.8 Flash Cyber, its most capable cybersecurity model, restricting access to vetted defenders through a new Fairwind Program rather than releasing it broadly, reflecting a growing industry consensus on cyber-capable AI risk.

Google shipped two new models this week from the same underlying foundation, and deliberately restricted access to one of them. Gemini 3.8 Flash is available to anyone, priced at 0.75 dollars per million input tokens and 3.75 dollars per million output tokens. Gemini 3.8 Flash Cyber, described by Google as its most capable cybersecurity model to date, is available only to a curated set of what the company calls trusted defenders, through a newly launched initiative called the Fairwind Program. The split is deliberate, and it says something significant about how seriously frontier labs now treat the dual-use risk baked into their own most capable systems.

What Flash Cyber can actually do

Flash Cyber is built specifically for autonomous vulnerability discovery, security research, and automated patching, and Google’s own published benchmarks back up the “most capable” claim with real numbers. On CyberGym, the standard industry benchmark for vulnerability discovery, Flash Cyber demonstrates what Google calls frontier-level performance, surpassing both its own predecessor, 3.5 Flash Cyber, and significantly larger frontier models. Google says it also cleared 70 percent or better on an internal benchmark testing vulnerability discovery across 20 different programming languages, a genuinely broad span for a single security-focused model to handle well, and posted 47.2 percent pass@1 on CWE-Bench, a benchmark tracking common weakness enumeration categories. Perhaps most concretely, Google claims Flash Cyber delivers 2.6 times more correct patches to real Chrome vulnerabilities than leading commercial alternatives, a specific, testable claim rather than a vague capability assertion.

The model is paired with CodeMender, Google’s own harness for validating and deploying fixes, letting defenders generate verified, deployment-ready patches in minutes inside their own secure cloud environment, according to Google’s own announcement, a dramatic compression compared to the weeks manual vulnerability remediation can typically take at enterprise scale.

Why access is deliberately restricted

The Fairwind Program already works with more than 650 partners globally, according to Google, and access is prioritized for four categories: governments and cyber authorities protecting public networks, critical infrastructure operators in healthcare, telecom, energy, and finance, maintainers of core technology platforms with wide-reaching software ecosystems, and approved cybersecurity teams conducting authorized defensive research. Applicants have to clear Google’s eligibility and due-diligence requirements, and organizations are expected to demonstrate a legitimate defensive use case while maintaining operational controls like multi-factor authentication and limited internal access before they’re granted entry.

Google’s own framing of the tradeoff is unusually direct for a product announcement: the same autonomous vulnerability-discovery capability that helps a defender patch a critical flaw before an attacker finds it could, in the wrong hands, help an attacker find that same flaw first. Restricting Flash Cyber to a vetted defender population is Google’s attempt to tilt that balance meaningfully toward defense, a real, structural bet that giving capable defenders a head start matters more than the theoretical efficiency loss from not releasing the model broadly.

Part of a broader industry pattern, not an isolated move

Google isn’t alone in reaching this conclusion. The Hacker News reported that Google, Anthropic, and OpenAI have each separately unveiled cyber-focused AI models, safeguards, and restricted-access programs around the same period, a pattern that reflects genuine, converging industry consensus rather than one company’s isolated caution. Anthropic has similarly held back its own cybersecurity-capable Mythos model under a program called Project Glasswing specifically because of its capacity to autonomously discover zero-day vulnerabilities, restricting access to a trusted coalition rather than releasing it broadly. This deliberate access-gating for cyber-capable models is also the direct product context behind the 117-company joint letter on AI cyber defense published just days before Flash Cyber’s release, in which OpenAI, Anthropic, Google, and more than a hundred other companies warned that AI-enabled cyberattacks will become significantly more widespread and sophisticated in the coming months.

Read together, these releases and that letter tell a consistent story: frontier labs increasingly believe the cyber-offense capability of their most advanced models has crossed a threshold serious enough to warrant genuinely restricted release, not just a safety disclaimer attached to an otherwise open product.

The general-purpose sibling tells its own story

Standard Gemini 3.8 Flash, released alongside Flash Cyber, is itself a genuinely capable general-purpose model, debuting at number 7 in Text Arena, ahead of Claude Opus 5, with real gains over its 3.7 Flash predecessor across multi-turn conversation, writing, coding, and business and financial reasoning tasks. On DeepSWE v1.1, a long-horizon software engineering benchmark, 3.8 Flash reportedly outperforms most larger frontier models at solving complex engineering problems end to end, at a fraction of the cost those larger models charge. That both models, the openly available Flash and the tightly restricted Flash Cyber, share the same foundational intelligence underscores that the restriction on Flash Cyber isn’t because Google lacks confidence in the underlying model’s quality. It’s a deliberate, calculated choice about who should have first access to a specific, high-risk capability.

What this signals about where the industry is heading

Google has been explicit that the Fairwind Program is a first step, not a finished framework, saying it will evolve access and product offerings alongside partner and user needs, and that it intends to collaborate with industry, governments, and the open-weight community to strike what it calls the right balance between open access and robust security. That language suggests the current restricted-access model isn’t necessarily permanent, but it also signals that Google doesn’t yet see a clear, safe path to opening Flash Cyber’s capabilities more broadly. For any organization not yet inside a program like Fairwind, Google’s public guidance points toward using CodeMender with its publicly available models on the Gemini Enterprise Agent Platform, combined with dedicated tools like AI Threat Defense, a meaningfully less capable but still genuinely useful path to some of the same defensive benefit.

See Google’s own Fairwind Program announcement for the complete eligibility criteria and technical detail.