NEWS
Microsoft’s 96% Cyber Score Sits at 86.3% on the Board
Microsoft posted a 96% CyberGym win for MAI-Cyber-1-Flash, then said the public leaderboard figure is 86.3% under a stricter scoring rule.
On August 13, Microsoft said the CyberGym figure on the public leaderboard for MDASH is 86.3%, not the 96% its executives posted in July. The 96% number is an any-crash score. The listed figure uses a stricter final-submission rule the benchmark added in July.
MAI-Cyber-1-Flash, the company’s first cybersecurity-specific model, still runs only inside MDASH for approved customers. Microsoft says that mix costs 50% less than its prior GPT stack. The model is not a public API.
Microsoft Posted Three CyberGym Scores, Not One
July launch posts treated 96% as a single win over Mythos, Gemini, and GPT. Seventeen days later, Mustafa Suleyman, CEO of Microsoft AI, and Hayete Gallot, executive vice president of Microsoft Security, added a clarification on CyberGym scores that splits the result into three different tests.
MICROSOFT’S THREE CYBERGYM FIGURES
| Scoring rule | Score | What it counts |
|---|---|---|
| Any-crash | 96% | An input that crashes the target, including bugs that are not the listed flaw |
| Target (any-of) | 90.4% | At least one candidate input maps to a known CyberGym vulnerability |
| Final-submission | 86.3% | The agent picks one input; Microsoft says this is the leaderboard number |
July materials also cited 95.95% for MDASH with MAI-Cyber-1-Flash and GPT-5.4. That figure sits next to the August any-crash label of 96% and should be read as the same looser metric, not a fourth test. The company says it filters edge cases that the evaluator might treat as valid crashes.
The gap is the product pitch. Executives compared 96% with a 12-point lead over Anthropic’s Mythos. The number Microsoft attributes to the public board, 86.3%, is below the 88.45% MDASH posted on May 12 under the older rule. A new scoring method, not a confessed miss, is the company’s account of that drop.
The July Pitch Was 96% at Half the Cost
On July 27, Suleyman put the rounded score, the Mythos gap, and the price cut in one line.
Big news! Our new MAI-Cyber-1-Flash model combined with MDASH, our multi agent security harness, delivers 96% on the CyberGym benchmark, 12pts above Mythos, at HALF the cost.
Proud of the team. More details in THREAD: pic.twitter.com/soH8AvvqHt
— Mustafa Suleyman (@mustafasuleyman) July 27, 2026
Satya Nadella, Microsoft’s chairman and CEO, wrote the same day that customers would get “frontier-grade security at half the cost.” Microsoft AI’s brand account said MDASH with MAI-Cyber-1-Flash and GPT-5.4 “out-performs Mythos by 12 points” and that the mix is “50% cheaper” than GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex.
Suleyman’s follow-up posts named the constraint he wants the specialist model to break. Token cost, he wrote, is now the limit for defenders who have to run agents all day. MAI-Cyber-1-Flash is built to take up to 90% of vulnerability detection and patching work so GPT-5.4 is reserved for the hardest 10%.
That routing story does not need a public leaderboard to be useful to a Defender customer. It does need a public board if the claim is a 12-point industry lead. Microsoft did not publish token counts, call volume, latency, or the task mix behind the 50% cut, so no one outside the company can replay the bill.
What CyberGym Level 1 Measures
CyberGym is a UC Berkeley benchmark. The public board ranks Level 1, where an agent gets a vulnerability write-up and the unpatched code, then must emit a working proof of concept. Success is a PoC that triggers on the pre-patch tree and fails on the patched tree. The suite holds 1,507 real-world vulnerability tasks drawn from 188 large projects that originally surfaced in OSS-Fuzz.
WHAT LEVEL 1 DOES NOT GRADE
- Blind discovery: The agent is handed a description of the bug, so the test is reproduction, not a cold hunt through unknown code.
- Patch quality: A crashing input can score even if the suggested fix is wrong or incomplete.
- Any crash: The board’s success rate is target reproduction with a working PoC, not a crash of any kind.
- A single model: Teams submit agents. Harnesses, extra memory, and extra trials are allowed and labelled.
Berkeley also warns that leading systems already sit in a high band, so small gaps may not mean much, and that runs are noisy. Microsoft’s any-crash score answers a wider question than the board asks. Its 86.3% final-submission figure is the one it maps onto that board.
Checked on July 28, the public list still showed Microsoft’s May entry rather than the new 95.95% launch number, with Wiz’s Atlas agent then listed first at 90.9%. Microsoft’s August note is the first time the company itself has said which of its figures belongs on that list.
A 5-Billion-Active Model Does Most of the Work
MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137 billion total parameters and 5 billion active, plus a 256,000-token context window. It is a cybersecurity fine-tune of MAI-Code-1-Flash, which itself comes from a MAI-Thinking-1 checkpoint. Text in, text out. Release date on the model card is July 27, 2026. Developer of record is Microsoft Ireland Operations Limited in Dublin.
THE ROUTING BET INSIDE MDASH
- Task share: The specialist is meant to cover up to 90% of MDASH work; GPT-5.4 takes the hardest 10%.
- Model swap: The evaluated setup replaced 80% of the prior MDASH models, a head-count change, not the same 90% task share.
- Bill: Microsoft says that mix is 50% cheaper than GPT-5.4 plus GPT-5.4 mini plus GPT-5.3 Codex.
- Access: Approved MDASH customers, Azure AI Foundry private preview, no standalone endpoint.
Those two percentages get mashed together in launch write-ups. They measure different things. One is how much work the small model is supposed to absorb. The other is how much of the old model roster got swapped out in the scored run.
Why the 50% Cost Cut Cannot Be Checked
Microsoft has not published the tokens, the number of model calls, the wait times, or the blend of easy and hard jobs behind that 50%. The comparison is only against its own last MDASH stack, not against Mythos or anyone else’s list price. Until those internals ship, the saving is a vendor claim, not a figure a buyer can audit.
The closed door is the other half of the same problem. MAI-Cyber-1-Flash has no public API, so a lab that wants to rerun CyberGym on the new weights cannot get them. The 96% line is a system result from a private, network-isolated setup. Microsoft says output can be wrong and should be reviewed before anyone acts on it.
The Model Card Still Shows Zeros on ExploitGym
Strip away MDASH and the picture changes. In a lightweight terminal harness, the same model is mediocre on several public security tests and blank on exploit generation.
STANDALONE SCORES ON THE MODEL CARD
| Test | Score | What the test asks |
|---|---|---|
| CVEBench | 0.314 | Exploit real web application flaws |
| CyberSecEval4 threat intel | 0.553 | Defensive threat-intelligence tasks |
| CyberSecEval4 malware analysis | 0.33 | Malware analysis tasks |
| CRSBench | 0.651 | Full-pipeline cyber reasoning, POV=1200 |
| ExploitGym kernel | 0 | Turn a bug into a working kernel exploit |
| ExploitGym userspace | 0 | Turn a bug into a userspace exploit |
| ExploitGym browser | 0 | Turn a bug into a browser exploit |
Those zeros match a defensive calibration. The card describes a model trained for discovery, triage, and patching inside MDASH, not for turning a crash into an exploit chain. A buyer who wants a 96% cyber model is not buying that standalone row. They are buying the harness around it.
The model is one input. The system is the product.
Taesoo Kim, Vice President of Agentic Security, Microsoft Security Blog, May 12, 2026
Kim’s line was already the MDASH thesis in May, before MAI-Cyber-1-Flash existed. The new weights are another input into that system. They are not a replacement for it.
MDASH Already Paid Off on a May Patch Tuesday
The harness is older than the cyber model. On May 12, Microsoft said MDASH had helped find 16 new Windows vulnerabilities in networking and authentication, including four Critical remote code execution bugs in tcpip.sys and the IKEv2 service. Those fixes went out in that day’s Patch Tuesday.
HOW THE HARNESS GOT HERE
- May 12, 2026: Microsoft publishes MDASH at 88.45% on CyberGym, then the top listed score, and reports 21 of 21 planted bugs found with zero false positives on StorageDrive, a private interview driver.
- July 27, 2026: MAI-Cyber-1-Flash is announced inside MDASH, with Project Perception named as the wider agent shell.
- August 3, 2026: Project Perception is scheduled to enter public preview inside Microsoft Defender.
- August 13, 2026: Microsoft posts the three-way CyberGym split and says 86.3% is the leaderboard figure.
MDASH is a pipeline, not a chatbot. Prepare builds an index and a threat model from past commits. Scan runs auditor agents. Validate runs debaters that argue reachability. Dedupe collapses repeats. Prove tries to build a triggering input. Microsoft says the harness holds more than 100 specialized agents and can swap models without throwing away plugins and scope files.
The Autonomous Code Security team built it. Several members came from Team Atlanta, which won DARPA’s AI Cyber Challenge and its $29.5 million prize. On internal history, Microsoft reported 96% recall against five years of confirmed MSRC cases in clfs.sys and 100% in tcpip.sys. That work predates MAI-Cyber-1-Flash. The new model is a cheaper engine in a machine that had already shipped bugs into Patch Tuesday.
Project Perception Opened Behind an Invite Wall
Project Perception is the name on the wider agent system. Gallot’s July 27 post frames it as a new cyber stack that coordinates three classes of specialized agents. Red team agents look for paths an attacker could take. Blue team agents decide which findings are real risk. Green team agents remediate and harden. MDASH’s vulnerability workflow is the first concrete job; Microsoft says Perception will later use MAI-Cyber-1-Flash on other security work.
The July event in San Francisco billed a public preview for August 3 inside Defender. Microsoft Learn now describes a Limited Public Preview for invited customers, for a defined window, before broader availability. Docs warn the product may change before a commercial release. In the Defender portal the surfaces are Overview, New Chat, Sessions, Agents, and Playbooks. Operators can approve, reject, or stop a run.
Microsoft also cites the data advantage it will not sell as a download: more than 100 trillion security signals a day and operational insight from 1.6 million customers, tied to MSRC cases and live defense across identity, endpoint, cloud, and apps. That loop is why the company argues a specialist model plus a harness beats one giant general model on a 24-hour security bill.
The public artifact from July is still a score. The thing a customer can actually get is an invited Defender preview of agents around a model they cannot call on their own. The 96% line remains an any-crash result from an isolated lab. The figure Microsoft maps to CyberGym’s board is 86.3%.
Frequently Asked Questions
What is the difference between MDASH and Project Perception?
MDASH is the multi-model scanning harness that prepares code, scans it, debates findings, drops duplicates, and tries to prove a bug with a triggering input. Project Perception is the broader Defender-side system that wraps red, blue, and green agents, playbooks, and a chat box around that kind of work and, later, other security jobs. MDASH can feed Perception; Perception is the product shell customers see in the portal.
Does CyberGym get harder if you hide the bug description?
Yes. Berkeley’s Level 0 setting, which withholds the text description of the target flaw, reproduces only 3.5% of instances. Level 1, the public board, hands the agent that write-up plus the unpatched tree. Success also falls as proof-of-concept inputs get longer: agents reach about 10% on cases whose ground-truth PoC is longer than 100 bytes, and those cases are 65.7% of the set.
Why does Microsoft keep GPT-5.4 in the scored mix?
MAI-Cyber-1-Flash is a 5-billion-active specialist, not a full replacement for a frontier model. Microsoft’s design is to let it handle the bulk of MDASH steps and to escalate the hardest tenth of tasks to GPT-5.4. The 96% and 86.3% figures both describe that two-model system, not the new weights alone.
Who can run MAI-Cyber-1-Flash today?
Microsoft’s model page says it is calibrated for defense and available only to verified defenders through MDASH. There is no public token price and no standalone Foundry endpoint for general callers. Sign-up is an interest form for MDASH, and Perception’s docs describe an invitation-only preview rather than an open download.
Did MDASH find bugs that were not already in CyberGym?
On Microsoft’s own code, yes. The May 12 write-up lists 16 Windows CVEs found with the harness, including CVE-2026-33827, a Critical unauthenticated use-after-free in tcpip.sys, and CVE-2026-33824, a Critical IKEv2 double-free that can lead to LocalSystem code execution. Those landed in the May 2026 Patch Tuesday cycle, separate from the public CyberGym corpus.
-
BUSINESS4 months agoMusk’s $914 Billion Lead Is a Public SpaceX Bet
-
SPORTS3 months agoFree Live Sports Streaming in 2026: What to Watch Without Cable
-
ENTERTAINMENT4 weeks agoSterling Point Holds No. 2 on Prime Video After 32 Days
-
ENTERTAINMENT3 months agoThe Odyssey’s IMAX Film Run Hit a 41-Theater Limit
-
ENTERTAINMENT3 months agoEndgame Encore’s $86 Million Trial for Infinity Vision
-
ENTERTAINMENT2 years agoAndrew Garfield’s Spider-Man Films Are No Longer Free
-
NEWS4 weeks agoAn AI Lung Drug Shifted Biological Age Clocks in Patients
-
GAMING4 weeks agoXbox Puts 15-Hour Caps on Game Pass Cloud Play
