Microsoft's AI system has identified critical Windows vulnerabilities
Microsoft has implemented a new AI system called MDASH to detect vulnerabilities in Windows. The system has already identified 16 issues, including four critical ones, before they could be exploited by malicious actors. This solution has demonstrated high effectiveness and is planned to be expanded for corporate clients.
Crius
Microsoft has implemented a new artificial intelligence system for detecting vulnerabilities in Windows, named MDASH (Multi-model Agentic Scanning Harness). This technology has already proven its effectiveness by identifying 16 vulnerabilities in the operating system before they could be exploited by malicious actors. Among the discovered issues, four were classified as critical, allowing remote code execution that could have led to unauthorized access to corporate networks. All identified vulnerabilities were addressed in the Patch Tuesday update on May 12.
How MDASH Works
MDASH was developed by Microsoft’s Autonomous Code Security team, which includes members of Team Atlanta—the winners of the DARPA AI Cyber Challenge. Unlike traditional scanners or single AI models, MDASH utilizes over 100 specialized agents based on various models. Each agent is responsible for a specific task: some search for vulnerabilities, others verify their authenticity, and in the final stage, the system attempts to generate input data to confirm the exploitability of the detected issue. Only after this process are the results forwarded to an engineer for further review.
Discovered Vulnerabilities
The system identified vulnerabilities in the following Windows components: the TCP/IP stack, IKEEXT IPsec service, HTTP.sys, Netlogon, Windows DNS, and the Telnet client. Ten of these were related to kernel mode, and most were accessible over the network without requiring authentication. Two of the four critical vulnerabilities stand out:
- CVE-2026-33827 — found in tcpip.sys, triggered by specially crafted IPv4 packets.
- CVE-2026-33824 — a double-free memory error in the IKEEXT service, accessible via UDP port 500 on devices with RRAS VPN, DirectAccess, or Always-On VPN enabled.
Both vulnerabilities allow attackers to gain LocalSystem privileges. Two other critical issues were found in Netlogon and Windows DNS Client, each receiving a CVSS score of 9.8.
Detecting some vulnerabilities, such as those in tcpip.sys, required analyzing three parallel code execution paths, all of which freed the same object. The IKEEXT issue involved six source files. Such multi-file and multi-threaded analysis is not possible with single-pass models.
Comparison with Other Solutions
MDASH achieved a score of 88.45% on the CyberGym UC Berkeley benchmark, which is based on 1,507 real-world vulnerability reproduction tasks, placing the system at the top of the public ranking. For comparison, Anthropic’s Mythos Preview model scored 83.1%, and OpenAI’s GPT-5.5 scored 81.8%. In closed testing on unpublished Windows StorageDrive drivers, MDASH detected all 21 injected vulnerabilities without any false positives. When tested on confirmed MSRC cases in clfs.sys and tcpip.sys over five years, the system demonstrated completeness rates of 96% and 100%, respectively.
MDASH is not tied to a specific model—Microsoft can update the models used as new ones become available, without having to redesign the entire pipeline. Currently, the system is available in a limited private mode for a small group of corporate clients, with plans to expand access in the coming months.
Market Trends and Threat Evolution
The announcement of MDASH follows similar initiatives from Anthropic’s Project Glasswing and OpenAI’s Daybreak, which are also developing comparable solutions with restricted access. All three companies aim to detect exploitable vulnerabilities before malicious actors do, and the gap between defense and attack capabilities powered by AI is rapidly closing.
At the same time, reports have emerged of the first known zero-day exploit created with the help of artificial intelligence and used in a large-scale attack to bypass two-factor authentication in a popular web server administration tool.
