Microsoft’s answer to AI-powered cyberattacks enters public preview today, August 3, 2026. Project Perception is a coordinated system of red, blue, and green AI agents built into Microsoft Defender, designed explicitly to defend against AI-speed attacks with AI-speed defense.
Quick facts
- Project Perception enters public preview August 3, 2026, following its unveiling at a Microsoft event on July 27.
- It coordinates three specialized agent types: red agents that probe systems like an attacker, blue agents that investigate like a security responder, and green agents that remediate and harden systems.
- The system runs on a new specialized model, MAI-Cyber-1-Flash, which Microsoft says scores 96% on the CyberGym benchmark, 12 points above the figure Microsoft cites for Anthropic’s Mythos.
- Pricing is consumption-based, measured in Security Compute Units (SCUs); Microsoft hasn’t published specific per-agent SCU rates.
- It’s the first major product from Microsoft’s security division under new security chief Hayete Gallot, who rejoined Microsoft from Google in February 2026.
How the three agent types actually work together
Per Microsoft’s own announcement, the design goal is closing the loop between finding a problem and fixing it without a manual hand-off at each step: a red agent identifies a vulnerability the way an attacker would, a blue agent investigates and prioritizes which flaws pose real risk, and a green agent writes and deploys the actual fix. Microsoft says the agents share intelligence through an orchestrated workflow and start with full organizational context, past incidents, identity relationships, and real-time signals, rather than working from a cold start on every task. Axios reports this mirrors the structure of a human security team, just running continuously rather than shift by shift.
The vulnerability-management numbers Microsoft is leading with
Microsoft’s first concrete use case is software vulnerability management, running MAI-Cyber-1-Flash inside its existing MDASH system, a multi-model team of agents already used internally. Microsoft says this configuration delivers 96% on the CyberGym benchmark, 12 points ahead of the score it attributes to Anthropic’s Mythos on the same test, at roughly half the cost of MDASH’s current production configuration. As with any benchmark figure a vendor publishes about its own product, treat that specific comparison as Microsoft’s own claim rather than independently verified until third-party evaluation catches up, the same caveat that applies to competing labs’ self-reported numbers.
Why Microsoft is building this alongside, not instead of, Security Copilot
Microsoft has been careful to position Project Perception as additive rather than a replacement for Security Copilot, its existing AI security assistant. Per Directions on Microsoft’s analysis, the company frames the distinction cleanly: Security Copilot is “AI that assists,” a generative chat interface a human still drives, while Perception is “AI that acts,” a system that takes autonomous action within defined boundaries. The two are meant to work together rather than compete for the same use case, with Perception initially scoped to Microsoft Defender before extending across the rest of Microsoft’s security product line over time.
Why this launch is happening now
The timing lines up directly with what the industry has been documenting all summer: CrowdStrike’s own August 3 threat report found AI-enabled attacks up 89% year-over-year, with exploitation windows now collapsing to within 24 to 48 hours of a vulnerability becoming public. Microsoft’s own pitch for Perception leans directly on that same argument: human-speed, alert-driven security operations can no longer keep pace with machine-speed attacks, and the only credible answer is defensive systems that also operate autonomously, not just faster human dashboards.
Key takeaway
Project Perception is a real, shipping product entering public preview today, not a roadmap promise, which puts it ahead of where most agentic security pitches currently sit. Its actual value will come down to how it performs against real, live threats at production scale, something no vendor’s own benchmark, including this one, can fully answer until independent security teams get their hands on it.


Leave a Reply