Microsoft introduces MAI Cyber 1 Flash through MDASH
Microsoft introduced MAI Cyber 1 Flash inside MDASH. Its reported CyberGym score covers a combined system, with access limited to verified defenders.

Correction added October 1, 2026. Microsoftโs official announcement is dated August 13, rather than the July 27 tracker date previously used here. It introduced MAI Cyber 1 Flash inside MDASH, a system for finding and remediating vulnerabilities.
The reported result measures a complete system, not the standalone model. That distinction matters when comparing both performance and cost.
CyberGym and cost claims
Microsoftโs roughly 96% CyberGym result belongs to the combined MDASH system using MAI Cyber 1 Flash and GPT 5.4. It uses any crash scoring. Microsoft separately reports 90.4% for target any of scoring and 86.3% for final submission scoring.
The reported 50% cost saving compares this combination with Microsoftโs prior MDASH configuration using GPT 5.4, GPT 5.4 mini, and GPT 5.3 Codex. It is not a general API price comparison.
MDASH integration
Microsoft describes MDASH as a system coordinating agents and models to identify, validate, and remediate vulnerabilities. MAI Cyber 1 Flash handles most tasks while larger models handle especially difficult work.
Microsoft describes tenant isolation, auditability, and sandboxed execution among MDASHโs controls. These measures do not remove the need for human review.
Access and safeguards
Benchmarks do not establish operational safety on their own. Deployment review should separately consider authorized access, monitoring, and how findings are validated.
Microsoftโs model page states that access is limited to verified defenders through MDASH. This is not an open weights release.
How it compares with GPT 5.6 Cyber
OpenAIโs GPT 5.6 Cyber is a separate model offered to approved defenders through Daybreak Red. The different systems and access paths should not be collapsed into a single benchmark comparison.
The comparison should stay tied to the named system, scoring method, and access conditions. A single benchmark percentage cannot replace that context.



