Microsoft MAI Cyber 1 Flash: What the 96% Score Measures
Microsoft launched MAI-Cyber-1-Flash, its first cybersecurity-specific model, inside MDASH, its multi-model vulnerability identification and remediation harness. Microsoft reports 95.95% on the CyberGym benchmark and says the configuration costs 50% less than its current best MDASH model mix.
Transcript
Microsoft just launched a cybersecurity model scoring near ninety six percent, so here is what that number measures.
Microsoft launched it inside MDASH, its vulnerability harness. But the score belongs to the whole system. The small model handles most tasks, GPT five point four the hardest.
CyberGym gives the agent a vulnerability description and the unpatched code, then checks for a working proof of concept. It does not test blind discovery or patch correctness.
Tested on its own, the model scored zero across ExploitGym's kernel, userspace, and browser categories. Microsoft says it may score higher inside the full system.
Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.
Sources
Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:
- Microsoft launched MAI-Cyber-1-Flash, its first cybersecurity-specific model, inside MDASH, its multi-model vulnerability identification and remediation harness, and reported 95.95 percent on the CyberGym benchmark.(Microsoft official announcement)
- The headline CyberGym result belongs to MDASH running MAI-Cyber-1-Flash alongside GPT-5.4, not to the new model on its own.(Microsoft official announcement)
- MAI-Cyber-1-Flash is designed to handle up to 90 percent of MDASH tasks, with GPT-5.4 reserved for the hardest 10 percent.(Microsoft official announcement)
- CyberGym Level 1 is a known-vulnerability reproduction test: it supplies a vulnerability description and the unpatched source code and checks whether the agent produces a working proof of concept. It does not measure blind vulnerability discovery and does not check whether a generated patch is correct.(The Hacker News (aggregator; benchmark description, not used in public attribution))
- The model card's External Cyber Benchmark Results table reports, verbatim, 'ExploitGym: Kernel = 0, Userspace = 0, Browser = 0'. Verified directly from the model card PDF on 2026-07-28.(MAI-Cyber-1-Flash model card)
- The model card states those standalone results come from evaluating MAI-Cyber-1-Flash as a single model on a lightweight terminal harness, and that 'when used within the codename MDASH system alongside other models, the model may support higher scores on these benchmarks'. Verified directly from the model card PDF on 2026-07-28.(MAI-Cyber-1-Flash model card)
- The evaluated configuration replaced 80 percent of MDASH's existing models and raised the reported CyberGym result from 88.4 percent to 95.95 percent.(MAI-Cyber-1-Flash model card)
