Connect with us

NEWS

Microsoft’s 96% Cyber Score Sits at 86.3% on the Board

Microsoft posted a 96% CyberGym win for MAI-Cyber-1-Flash, then said the public leaderboard figure is 86.3% under a stricter scoring rule.

Published

on

On August 13, Microsoft said the CyberGym figure on the public leaderboard for MDASH is 86.3%, not the 96% its executives posted in July. The 96% number is an any-crash score. The listed figure uses a stricter final-submission rule the benchmark added in July.

MAI-Cyber-1-Flash, the company’s first cybersecurity-specific model, still runs only inside MDASH for approved customers. Microsoft says that mix costs 50% less than its prior GPT stack. The model is not a public API.

Microsoft Posted Three CyberGym Scores, Not One

July launch posts treated 96% as a single win over Mythos, Gemini, and GPT. Seventeen days later, Mustafa Suleyman, CEO of Microsoft AI, and Hayete Gallot, executive vice president of Microsoft Security, added a clarification on CyberGym scores that splits the result into three different tests.

MICROSOFT’S THREE CYBERGYM FIGURES

Scoring rule Score What it counts
Any-crash 96% An input that crashes the target, including bugs that are not the listed flaw
Target (any-of) 90.4% At least one candidate input maps to a known CyberGym vulnerability
Final-submission 86.3% The agent picks one input; Microsoft says this is the leaderboard number

July materials also cited 95.95% for MDASH with MAI-Cyber-1-Flash and GPT-5.4. That figure sits next to the August any-crash label of 96% and should be read as the same looser metric, not a fourth test. The company says it filters edge cases that the evaluator might treat as valid crashes.

The gap is the product pitch. Executives compared 96% with a 12-point lead over Anthropic’s Mythos. The number Microsoft attributes to the public board, 86.3%, is below the 88.45% MDASH posted on May 12 under the older rule. A new scoring method, not a confessed miss, is the company’s account of that drop.

The July Pitch Was 96% at Half the Cost

On July 27, Suleyman put the rounded score, the Mythos gap, and the price cut in one line.

Satya Nadella, Microsoft’s chairman and CEO, wrote the same day that customers would get “frontier-grade security at half the cost.” Microsoft AI’s brand account said MDASH with MAI-Cyber-1-Flash and GPT-5.4 “out-performs Mythos by 12 points” and that the mix is “50% cheaper” than GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex.

Suleyman’s follow-up posts named the constraint he wants the specialist model to break. Token cost, he wrote, is now the limit for defenders who have to run agents all day. MAI-Cyber-1-Flash is built to take up to 90% of vulnerability detection and patching work so GPT-5.4 is reserved for the hardest 10%.

That routing story does not need a public leaderboard to be useful to a Defender customer. It does need a public board if the claim is a 12-point industry lead. Microsoft did not publish token counts, call volume, latency, or the task mix behind the 50% cut, so no one outside the company can replay the bill.

What CyberGym Level 1 Measures

CyberGym is a UC Berkeley benchmark. The public board ranks Level 1, where an agent gets a vulnerability write-up and the unpatched code, then must emit a working proof of concept. Success is a PoC that triggers on the pre-patch tree and fails on the patched tree. The suite holds 1,507 real-world vulnerability tasks drawn from 188 large projects that originally surfaced in OSS-Fuzz.

WHAT LEVEL 1 DOES NOT GRADE

  • Blind discovery: The agent is handed a description of the bug, so the test is reproduction, not a cold hunt through unknown code.
  • Patch quality: A crashing input can score even if the suggested fix is wrong or incomplete.
  • Any crash: The board’s success rate is target reproduction with a working PoC, not a crash of any kind.
  • A single model: Teams submit agents. Harnesses, extra memory, and extra trials are allowed and labelled.

Berkeley also warns that leading systems already sit in a high band, so small gaps may not mean much, and that runs are noisy. Microsoft’s any-crash score answers a wider question than the board asks. Its 86.3% final-submission figure is the one it maps onto that board.

Checked on July 28, the public list still showed Microsoft’s May entry rather than the new 95.95% launch number, with Wiz’s Atlas agent then listed first at 90.9%. Microsoft’s August note is the first time the company itself has said which of its figures belongs on that list.

A 5-Billion-Active Model Does Most of the Work

MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137 billion total parameters and 5 billion active, plus a 256,000-token context window. It is a cybersecurity fine-tune of MAI-Code-1-Flash, which itself comes from a MAI-Thinking-1 checkpoint. Text in, text out. Release date on the model card is July 27, 2026. Developer of record is Microsoft Ireland Operations Limited in Dublin.

THE ROUTING BET INSIDE MDASH

  • Task share: The specialist is meant to cover up to 90% of MDASH work; GPT-5.4 takes the hardest 10%.
  • Model swap: The evaluated setup replaced 80% of the prior MDASH models, a head-count change, not the same 90% task share.
  • Bill: Microsoft says that mix is 50% cheaper than GPT-5.4 plus GPT-5.4 mini plus GPT-5.3 Codex.
  • Access: Approved MDASH customers, Azure AI Foundry private preview, no standalone endpoint.

Those two percentages get mashed together in launch write-ups. They measure different things. One is how much work the small model is supposed to absorb. The other is how much of the old model roster got swapped out in the scored run.

Why the 50% Cost Cut Cannot Be Checked

Microsoft has not published the tokens, the number of model calls, the wait times, or the blend of easy and hard jobs behind that 50%. The comparison is only against its own last MDASH stack, not against Mythos or anyone else’s list price. Until those internals ship, the saving is a vendor claim, not a figure a buyer can audit.

The closed door is the other half of the same problem. MAI-Cyber-1-Flash has no public API, so a lab that wants to rerun CyberGym on the new weights cannot get them. The 96% line is a system result from a private, network-isolated setup. Microsoft says output can be wrong and should be reviewed before anyone acts on it.

The Model Card Still Shows Zeros on ExploitGym

Strip away MDASH and the picture changes. In a lightweight terminal harness, the same model is mediocre on several public security tests and blank on exploit generation.

STANDALONE SCORES ON THE MODEL CARD

Test Score What the test asks
CVEBench 0.314 Exploit real web application flaws
CyberSecEval4 threat intel 0.553 Defensive threat-intelligence tasks
CyberSecEval4 malware analysis 0.33 Malware analysis tasks
CRSBench 0.651 Full-pipeline cyber reasoning, POV=1200
ExploitGym kernel 0 Turn a bug into a working kernel exploit
ExploitGym userspace 0 Turn a bug into a userspace exploit
ExploitGym browser 0 Turn a bug into a browser exploit

Those zeros match a defensive calibration. The card describes a model trained for discovery, triage, and patching inside MDASH, not for turning a crash into an exploit chain. A buyer who wants a 96% cyber model is not buying that standalone row. They are buying the harness around it.

The model is one input. The system is the product.

Taesoo Kim, Vice President of Agentic Security, Microsoft Security Blog, May 12, 2026

Kim’s line was already the MDASH thesis in May, before MAI-Cyber-1-Flash existed. The new weights are another input into that system. They are not a replacement for it.

MDASH Already Paid Off on a May Patch Tuesday

The harness is older than the cyber model. On May 12, Microsoft said MDASH had helped find 16 new Windows vulnerabilities in networking and authentication, including four Critical remote code execution bugs in tcpip.sys and the IKEv2 service. Those fixes went out in that day’s Patch Tuesday.

HOW THE HARNESS GOT HERE

  1. May 12, 2026: Microsoft publishes MDASH at 88.45% on CyberGym, then the top listed score, and reports 21 of 21 planted bugs found with zero false positives on StorageDrive, a private interview driver.
  2. July 27, 2026: MAI-Cyber-1-Flash is announced inside MDASH, with Project Perception named as the wider agent shell.
  3. August 3, 2026: Project Perception is scheduled to enter public preview inside Microsoft Defender.
  4. August 13, 2026: Microsoft posts the three-way CyberGym split and says 86.3% is the leaderboard figure.

MDASH is a pipeline, not a chatbot. Prepare builds an index and a threat model from past commits. Scan runs auditor agents. Validate runs debaters that argue reachability. Dedupe collapses repeats. Prove tries to build a triggering input. Microsoft says the harness holds more than 100 specialized agents and can swap models without throwing away plugins and scope files.

The Autonomous Code Security team built it. Several members came from Team Atlanta, which won DARPA’s AI Cyber Challenge and its $29.5 million prize. On internal history, Microsoft reported 96% recall against five years of confirmed MSRC cases in clfs.sys and 100% in tcpip.sys. That work predates MAI-Cyber-1-Flash. The new model is a cheaper engine in a machine that had already shipped bugs into Patch Tuesday.

Project Perception Opened Behind an Invite Wall

Project Perception is the name on the wider agent system. Gallot’s July 27 post frames it as a new cyber stack that coordinates three classes of specialized agents. Red team agents look for paths an attacker could take. Blue team agents decide which findings are real risk. Green team agents remediate and harden. MDASH’s vulnerability workflow is the first concrete job; Microsoft says Perception will later use MAI-Cyber-1-Flash on other security work.

The July event in San Francisco billed a public preview for August 3 inside Defender. Microsoft Learn now describes a Limited Public Preview for invited customers, for a defined window, before broader availability. Docs warn the product may change before a commercial release. In the Defender portal the surfaces are Overview, New Chat, Sessions, Agents, and Playbooks. Operators can approve, reject, or stop a run.

Microsoft also cites the data advantage it will not sell as a download: more than 100 trillion security signals a day and operational insight from 1.6 million customers, tied to MSRC cases and live defense across identity, endpoint, cloud, and apps. That loop is why the company argues a specialist model plus a harness beats one giant general model on a 24-hour security bill.

The public artifact from July is still a score. The thing a customer can actually get is an invited Defender preview of agents around a model they cannot call on their own. The 96% line remains an any-crash result from an isolated lab. The figure Microsoft maps to CyberGym’s board is 86.3%.

Frequently Asked Questions

What is the difference between MDASH and Project Perception?

MDASH is the multi-model scanning harness that prepares code, scans it, debates findings, drops duplicates, and tries to prove a bug with a triggering input. Project Perception is the broader Defender-side system that wraps red, blue, and green agents, playbooks, and a chat box around that kind of work and, later, other security jobs. MDASH can feed Perception; Perception is the product shell customers see in the portal.

Does CyberGym get harder if you hide the bug description?

Yes. Berkeley’s Level 0 setting, which withholds the text description of the target flaw, reproduces only 3.5% of instances. Level 1, the public board, hands the agent that write-up plus the unpatched tree. Success also falls as proof-of-concept inputs get longer: agents reach about 10% on cases whose ground-truth PoC is longer than 100 bytes, and those cases are 65.7% of the set.

Why does Microsoft keep GPT-5.4 in the scored mix?

MAI-Cyber-1-Flash is a 5-billion-active specialist, not a full replacement for a frontier model. Microsoft’s design is to let it handle the bulk of MDASH steps and to escalate the hardest tenth of tasks to GPT-5.4. The 96% and 86.3% figures both describe that two-model system, not the new weights alone.

Who can run MAI-Cyber-1-Flash today?

Microsoft’s model page says it is calibrated for defense and available only to verified defenders through MDASH. There is no public token price and no standalone Foundry endpoint for general callers. Sign-up is an interest form for MDASH, and Perception’s docs describe an invitation-only preview rather than an open download.

Did MDASH find bugs that were not already in CyberGym?

On Microsoft’s own code, yes. The May 12 write-up lists 16 Windows CVEs found with the harness, including CVE-2026-33827, a Critical unauthenticated use-after-free in tcpip.sys, and CVE-2026-33824, a Critical IKEv2 double-free that can lead to LocalSystem code execution. Those landed in the May 2026 Patch Tuesday cycle, separate from the public CyberGym corpus.

Harry is the editor and publisher of MIND CRON, an independent title built on ten years of journalism that took him from the reporter's notebook to the editor's chair. Breaking news is where his rules are strictest. A story goes out when the primary document is in hand or two independent sources confirm the same fact, and not before, however loud the rumour. Anything still moving is labelled as developing, each update carries the time it was made, and the original wording stays visible so readers can see what changed. That discipline applies whether the story is a market shock in business, an outage in technology, a result in sports, a launch in gaming or a recall in auto, and it is no looser for science, entertainment, lifestyle, travel or the wider news pages. Numbers are checked against the source before publication. Errors are corrected openly under a public corrections policy. Tips from readers are checked the same way as everything else, and Harry reads and answers that mail himself at support@mindcron.com.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending