Microsoft has launched its first cybersecurity-specific model inside MDASH, its multi-model vulnerability identification and remediation harness.

The company says MDASH, using MAI-Cyber-1-Flash and GPT-5.4, scored 95.95% on CyberGym. It also claims the configuration costs 50% less than its current best MDASH combination of GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. Access is limited to approved MDASH customers through an Azure AI Foundry private preview.

MAI-Cyber-1-Flash is designed to handle up to 90% of MDASH tasks, with GPT-5.4 reserved for the hardest 10%. It is available only inside MDASH, not as a standalone public model or general-purpose application programming interface.

The headline score belongs to MDASH running MAI-Cyber-1-Flash alongside GPT-5.4, not to the new model by itself. CyberGym Level 1 is a known-vulnerability reproduction test. It gives an agent a vulnerability description and the corresponding unpatched source code, then checks whether it can produce a working proof of concept. It does not measure blind vulnerability discovery or whether a generated patch is correct.

CyberGym's public leaderboard did not list Microsoft's 95.95% result when checked on July 28, 2026. It listed Wiz's Atlas agent first at 90.9% in a July 27 entry, while Microsoft's May 12 MDASH submission remained at 88.4%. Microsoft's public materials do not say whether the result was submitted for listing.

Microsoft's earlier 96.55% MDASH result does not resolve the comparison. That June figure counted any crash, including non-target vulnerabilities. The July materials do not say whether the 95.95% result uses the same criterion, so the two scores cannot safely be read as a before-and-after performance trend.

According to Microsoft's model card, MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137 billion total parameters, five billion active parameters, and a 256,000-token context window. It is a cybersecurity fine-tune of MAI-Code-1-Flash, which was developed from a MAI-Thinking-1 mid-training checkpoint.

The model card says the evaluated configuration replaced 80% of MDASH's existing models and raised the reported CyberGym result from 88.4% to 95.95%. That 80% figure is the share of models replaced. The separate 90% figure is the maximum share of tasks Microsoft says the smaller model can handle.

Taken together, the disclosed design points to routing as the central technical claim: MAI-Cyber-1-Flash is intended to handle most tasks, GPT-5.4 takes the hardest remainder, and Microsoft reports the outcome at the MDASH system level.

Microsoft's launch announcement defines the 50% saving against its current best MDASH model mix of GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. The product page separately describes the system as delivering "comparable performance at 50% of the cost of leading models." The announcement and model card do not disclose the token use, call volume, latency, task mix, or compute allocation behind that comparison, so the figure cannot yet be independently reproduced or normalised against other systems.

"The model is one input, the system around it is the product."

Taesoo Kim, Microsoft's vice president of agentic security, used that distinction when describing MDASH in June. Under a lightweight terminal harness, the model card reports scores of 0.314 on CVEBench, 0.553 on CyberSecEval4 threat intelligence, 0.33 on its malware-analysis test, and 0.651 on CRSBench at POV=1200.

The model scored zero across the kernel, userspace, and browser categories of ExploitGym, which asks agents to turn supplied vulnerabilities and crashing inputs into working code-execution exploits. Those results come from different tasks and scoring scales, so none is a standalone CyberGym score for MAI-Cyber-1-Flash.

Microsoft said all benchmark testing took place in a network-isolated environment with no access to production systems, the public internet, or external services. The model card also warns that generated text and code may be inaccurate or incomplete and should be reviewed before consequential use.

Software vulnerability management using MAI-Cyber-1-Flash inside MDASH is the first scenario Microsoft has announced for Project Perception, its broader system for coordinating defensive security agents. Project Perception is scheduled to enter public preview on August 3, with Microsoft planning to extend the model beyond software vulnerability work to additional security workflows.

Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.