The Agentic Post
Breaking
Digital Twins and Physical AI  ·  Humanoid Robots in Manufacturing  ·  AI Data Centers and Water Usage  ·  The AI Chip Supply Chain, Explained  ·  What Is Fine-Tuning? A Plain Explainer  ·  Meta and Sierra Want to Give AI Agents a Front Door to Stores  ·  
Home/AI Safety
Anthropic Opens Mythos to More Defenders in Three Tiers

Anthropic Opens Mythos to More Defenders in Three Tiers

AI Safety

Anthropic merged Project Glasswing into an expanded Cyber Verification Program with three access tiers, and said Glasswing partners found at least 129,000 verified vulnerabilities between April and July.

Anthropic expanded its Cyber Verification Program on October 6, 2026, merging it with Project Glasswing into a single programme with three access tiers. Each tier gives vetted security professionals fewer cyber blocks on Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1. Anthropic also disclosed that Glasswing partners found at least 129,000 verified software vulnerabilities between April and July 2026, more than 33,000 of them rated critical or high severity.

Why do security teams need a special programme at all?

Because Anthropic’s public models deliberately block most cyber work. Cybersecurity is dual use: the steps that let a defender prove and patch a flaw are the same steps an attacker uses to exploit it. So the generally available versions of Opus 5.5, Fable 5.1 and Sonnet 5.5 refuse a lot of legitimate security tasks. The programme is how verified defenders get those blocks lifted, in proportion to how much they need.

What are the three tiers?

  • Defense Access: incident response, malware reverse-engineering, and validating vulnerabilities. Open to in-house security teams, critical infrastructure operators of any size such as regional hospitals or local utilities, smaller security firms, open-source maintainers, and individual researchers with a track record. Anthropic aims to answer applications within a few days.
  • Red Team Access: adds authorised penetration testing. Organisations only, review takes a few weeks, and applicants sit in Defense Access meanwhile. Actions that could cause physical harm or mass disruption, such as deploying ransomware, stay blocked.
  • Specialized Access: fewest blocks, for a small set of organisations authorised to test systems where failure could hurt people or disrupt markets, such as flight software, power grids, telecoms and interbank payments. Reviewed in collaboration with the US government. Existing Glasswing members move here automatically.

Do the tiers actually work?

Anthropic tested Opus 5.5 on CyScenarioBench, ten multi-stage cyber operation challenges, five attempts each per tier. Without the programme, every task was blocked on the first prompt. In Defense Access, 46 of 50 attempts were blocked at some point. In Red Team Access, nothing was blocked and Opus completed 34 of 50, matching its 67.6% success rate with no safeguards. That is Anthropic grading its own controls, but the spread between tiers is the behaviour you would want.

How credible are the 129,000 vulnerabilities?

Treat it as a lower bound with soft edges. The figure comes from survey data from 33 partner reports, Anthropic says the real number is likely at least five times higher, and fewer than half of partners disclosed how many flaws they had patched. Anthropic’s own open-source scanning found another 5,500 between April and October. The scale is plausible, and partners such as Booz Allen and Comcast have described months or years of work compressed, but the methodology is self-reported.

Is there a privacy catch?

Yes. Enrolled organisations must allow data retention so Anthropic can monitor for misuse, which is a real cost for teams handling sensitive incidents. Anthropic says a feature called Enterprise Frontier Safeguards, due later this fall, will let eligible organisations keep that data in cloud infrastructure they control. The programme runs on the Claude Platform, Google Cloud Vertex AI and Microsoft Foundry, but on Amazon Bedrock only for customers eligible for that new feature.

The timing is pointed. Mistral launched Large 4 the same day, highlighting that Opus 5.5 scores near zero on one vulnerability test because it refuses. This is Anthropic’s answer: the capability exists, gated by verification rather than open weights. It extends the same logic as Google’s Fairwind programme, and lands a week after Microsoft warned that attackers now weaponise flaws in under 24 hours.

Apply or read the details in Anthropic’s announcement.

Up Next
Mistral Large 4 Is a 1-Trillion-Parameter Open Model Built in Europe

Mistral Large 4 Is a 1-Trillion-Parameter Open Model Built in Europe

Claude

Mistral previewed Large 4, a 1-trillion-parameter multimodal open-weight model trained on 3,800 GPUs in Europe, with weights due end of October and top-five cybersecurity scores that closed models refuse to match.

Mistral launched a public preview of Mistral Large 4 on October 6, 2026: a 1-trillion-parameter, natively multimodal model with 49 billion parameters active per token. It is Mistral’s largest model, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European data centres. Mistral says the open weights will be released by the end of October. The preview API costs 1.36 dollars per million input tokens and 4.18 per million output. The company’s own nickname for it is "le Chonk."

How good is it?

Strong for an open model, behind the closed frontier. On coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4.0, for a combined Coding Agent Index of 49.8% that puts it ahead of DeepSeek V4 Pro and Qwen3.8-Max. On AutomationBench, 657 business workflows across tools like Gmail and Salesforce, it scores 59.9%.

For context, Claude Opus 5.5 scores 66.4% on Terminal-Bench 4.0 against Large 4’s 28.3%. In a blind human evaluation of coding quality, Large 4 ranked second of five behind Claude Opus 5. Its preview scored 38 on the Artificial Analysis Intelligence Index, the best Western open-weight result but still behind the leading Chinese open models.

What is the cybersecurity claim?

This is the part worth reading carefully. Mistral says Large 4 ranks in the top five on the Artificial Analysis Cyber Index and scores 82% on a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, the highest of any model. It solves 93% of the 40 Cybench challenges.

Mistral also points out that Claude Opus 5.5 and GPT-6 Astra score near zero on that same reproduce-and-patch test because they refuse the task. That is a genuine point about defenders being blocked mid-incident. It is also exactly the capability the closed labs have deliberately restricted, through programmes like Google’s Fairwind and Anthropic’s cyber verification tiers. Shipping it as open weights means anyone can run it without a vetting step, which cuts both ways.

Mistral says it is red-teaming the model with cybersecurity firms, vetted partners and state authorities before the weights go out, and that Large 4 refuses malicious cyber prompts more often than other open models on JailbreakBench, StrongREJECT and AgentHarm.

Why does "trained in Europe" matter?

Sovereignty. Mistral trained and serves the model on its own infrastructure and will offer a European deployment it runs end-to-end under European law, independent of US cloud providers. For European governments and regulated firms that cannot send data to American services, that is the selling point more than any benchmark. The model is the first output of Mistral’s 3 billion euro Series D, which it calls the largest equity round ever raised by a European tech company.

Should you use it?

If you need open weights, European data residency or security research without refusals, it is now the strongest Western option. If you need the best coding agent available, the closed frontier models are still well ahead on Terminal-Bench. And the usual caveat applies: most numbers above are Mistral’s own or privately run evaluations, and the reinforcement learning run is still in progress, so the released weights may score differently. It continues the open-model squeeze on pricing we covered with AT&T moving traffic to open models.

See Mistral’s announcement.