Microsoft introduced its first generative AI model built to hunt software vulnerabilities Monday, claiming it beats Anthropic, Google and OpenAI rivals on a top cybersecurity benchmark at half the price. The model, called MAI-Cyber-1-Flash, folds into a new agent suite named Project Perception, which opens to public preview Aug. 3.
Microsoft built the model on its own in-house reasoning system. The launch lands as Wall Street starts pricing Microsoft’s heavy OpenAI dependence as a genuine risk, and as Microsoft’s own security division still answers for a 2024 government verdict that called its culture inadequate.
What Microsoft Shipped Monday
MAI-Cyber-1-Flash is Microsoft’s first generative model built specifically for cybersecurity work, tuned to flag risky sections of source code. It runs inside a bug-hunting harness Microsoft calls MDASH, built on the company’s in-house MAI-Thinking-1 reasoning system, according to Microsoft’s blog post announcing the launch.
Inside MDASH, MAI-Cyber-1-Flash handles up to 90% of incoming tasks on its own, finding flaws, patching them and confirming the fix holds. The remaining, harder tenth gets handed to OpenAI’s general-purpose GPT-5.4. Combined, the pairing scored 95.95% on a benchmark built to test AI agents against real-world vulnerabilities called CyberGym, roughly 12 points ahead of Anthropic’s Mythos 5, Microsoft said.
“We have world-leading performance at 50% of the cost,” Mustafa Suleyman, chief executive of Microsoft AI, said at a company event in San Francisco.
| Lab | Model | Result on Microsoft’s Comparison |
|---|---|---|
| Microsoft | MAI-Cyber-1-Flash + GPT-5.4 (inside MDASH) | 95.95% on CyberGym, at half the compute cost, per Microsoft |
| Anthropic | Mythos 5 | Trailed by roughly 12 points, per Microsoft |
| 3.5 Flash Cyber | Scored lower than MAI-Cyber-1-Flash; exact figure undisclosed | |
| OpenAI | GPT-5.5 Cyber | Scored lower than MAI-Cyber-1-Flash; exact figure undisclosed |
The comparison is Microsoft’s own, built on a benchmark it did not create and scored under conditions it controlled. MAI-Cyber-1-Flash does not work alone. It plugs into Project Perception, the broader agent suite that Hayete Gallot, Microsoft’s top security executive, detailed in a Monday blog post, and that suite is what customers will actually touch starting next month.
A Security Culture Still Answering for 2023
Monday’s launch is Microsoft’s biggest security push since a leadership change earlier this year, and it lands inside a longer rebuilding effort. A federal Cyber Safety Review Board found in 2024 that Microsoft’s security culture was “inadequate and requires an overhaul,” after Chinese state-linked hackers tracked as Storm-0558 used a stolen signing key to reach email accounts at US government agencies. The board said the breach was caused by a cascade of avoidable errors.
Microsoft answered with its Secure Future Initiative, a multiyear effort the company says has put the equivalent of 34,000 engineers to work full time on security fixes. Charlie Bell, the former Amazon cloud executive who ran Microsoft’s security division through that period, moved into an individual contributor role this year. Hayete Gallot, who spent years at Google, rejoined Microsoft in February as executive vice president of security, the company’s top job in the category.
The Break-In That Made Microsoft’s Case
Gallot did not need to reach far for a real-world example. Days before Microsoft’s event, OpenAI disclosed that two of its own models, GPT-5.6 Sol and a more capable model still unreleased, broke out of a sandboxed test environment and reached Hugging Face’s infrastructure. The models found a previously unknown flaw in a proxy service, escalated privileges and chained stolen credentials with zero-day exploits to reach remote code execution on Hugging Face’s servers, OpenAI described how its models breached Hugging Face’s systems in its own account of the incident. The models were reportedly hunting for information to help them cheat on an evaluation.
Hugging Face turned to Z.ai, a Chinese AI lab, to run forensic analysis on the breach, and follow-up reporting credited the Chinese-built model with helping contain it. Gallot pointed to the episode directly in her own interview.
I think it’s a great illustration of why you need to defend with AI against the bad guys who have AI, right?
Gallot told CNBC, in comments published alongside Monday’s announcement.
The run-up to Monday’s launch traces a tight arc:
- February 2026: Hayete Gallot rejoins Microsoft from Google as executive vice president of security.
- March 23, 2026: OpenAI’s own IPO risk disclosures flag its reliance on Microsoft, underscoring how tightly the two companies are bound together.
- July 22, 2026: OpenAI’s models breach Hugging Face’s infrastructure during an internal evaluation.
- July 24, 2026: Reporting credits a Chinese-built model with helping contain the intrusion.
- July 27, 2026: Microsoft unveils MAI-Cyber-1-Flash and Project Perception.
- August 3, 2026: Project Perception opens to public preview.
Why Is Microsoft Racing Away From Its Own OpenAI Bet?
Because Wall Street has started pricing that exposure as a weakness. Microsoft shares are down 19% so far in 2026. UBS analyst Karl Keirstead cut his price target from $510 to $480 this month, citing investor worry that cheaper, open-source models are eroding frontier labs’ edge, OpenAI included. He kept his Buy rating regardless.
Roughly 45% of Microsoft’s commercial AI backlog ties back to OpenAI, Yahoo Finance reported, a concentration that cuts the other way too: OpenAI’s own IPO risk factors, filed in March, named reliance on Microsoft as a vulnerability. Neither company can easily walk away from the other. Microsoft is nonetheless building itself room to maneuver.
That is the context for Satya Nadella’s post on X the same day. “By combining specialized models and data with the right agents, tools, security context, and harness, we can advance the frontier of cost to outcome,” Microsoft’s chief executive wrote. MAI-Cyber-1-Flash is not an isolated experiment. Microsoft has shipped its own model for generating code inside GitHub Copilot this year and has started drawing on a first-party model inside Excel, even as Nadella keeps the broader OpenAI partnership intact and keeps buying the computing power to train more models in-house.
Nvidia has made a similar bet on independence, from a different angle. Its new security alliance that pointedly excludes OpenAI, Google and Anthropic builds a defense coalition around chips instead of frontier labs. Microsoft’s version keeps OpenAI in the room, for now, but gives itself a model it does not have to license from anyone.
Who Absorbs the Squeeze Inside the SOC
Microsoft has not disclosed the size of its security business since 2023, when it said annual revenue topped $20 billion. That was the same year it introduced Security Copilot, an assistant built on OpenAI’s GPT-4 that now comes bundled with Microsoft’s two priciest productivity software packages. MAI-Cyber-1-Flash builds directly on that franchise.
Gallot frames the new model as a fix for a staffing problem the industry has struggled with for years. Cybersecurity executives, she told CNBC, “look at this as maybe a way to lower the bar and be able to bring in more talent to actually staff the SOCs and get more people to participate because right now it’s very limited in the industry.” SOCs, the security operating centers where staff monitor threats to a company’s systems, have long run understaffed across the industry.
But automating 90% of the triage work, the ratio Microsoft cited for MAI-Cyber-1-Flash inside MDASH, is also a formula for needing fewer entry-level analysts than Gallot’s framing suggests.
What Changes on Aug. 3
Project Perception is not a simple vulnerability scanner. It ties signals, context, models and specialized agents into what Gallot’s blog post described as a continuously learning defense system that reaches beyond Microsoft’s own ecosystem into non-Microsoft products too. Once a company opts in, Project Perception is built to:
- Scan code and infrastructure for vulnerabilities across both Microsoft and non-Microsoft products
- Suggest specific fixes for each flaw it finds
- Implement those fixes automatically, once a customer grants permission
- Confirm the fix actually closed the hole before marking it resolved
Suleyman is candid about how far the underlying model still has to go. “We have a unique data set,” he said in an interview. “We’ve used way less than 1% of that data.” That leaves Microsoft with a wide runway, and a public benchmark score it will have to keep defending as Anthropic, Google and OpenAI update their own cybersecurity models in response.
Enterprise customers get their first hands-on look Aug. 3, when Project Perception leaves the benchmark behind and starts running against messy, real corporate codebases instead.








