AI models for the first time received their own "cyberweapons index," which evaluates not knowledge of the hacking, but the ability to independently conduct a real chain of attack. In the first Cyber Weapon Index, Booz Allen tested 18 American and Chinese models, and the result was extremely uneven. Claude Mythos scored 80 points and became the only model that was able to autonomously pass the entire chain from intelligence to complete domain capture.
The gap between the first and second place was huge. Claude Mythos received 80 points, Grok-4.5 scored 49, GPT-5.6 Sol 46, Muse Spark 1.1 and Chinese Kimi K3 at 38, and GLM-5.2 from Z.ai 37. At the bottom of the ranking were Claude Sonnet 5 with 13 points, GLM-4.5-Air with 11, Qwen3.6-35B with 9 and Qwen3-Coder with 4. Booz Allen believes the backlog will shorten rapidly and most of the models tested will be able to approach Mythos capabilities within about six months. The latest forecast reflects the assessment of researchers, rather than the confirmed pace of model development.
The name Cyber Weapon Index sounds loud, but the methodology evaluates quite specific skills. The final score is equal to two indicators. Vulnerability Research Score shows how well the model self-founds vulnerabilities in compiled code without source. Kill Chain Attainment Score measures how far the model will advance along the attack chain in the secure enterprise Active Directory network. The formula is extremely simple, CWI is equal to the average value of two estimates.
For the second part of the test, the researchers gave each model a separate attacking machine and allowed to independently execute commands against the working test corporate network. The models were explored, gained initial access, fixed, mined credentials, increased privileges, performed lateral movement and tried to gain full control over the domain. The tests were carried out both with and without pre-provided stolen credentials. The result was verified by network traffic, host logs, domain controllers, and intrusion detection systems, so models could not get points for a simple statement about successful hacking.
Fraudsters learn every day. You are not.
SecurityLab in Telegram.
Claude Mythos has noticeably stood out on the practical part. With the current credentials, the model in all attempts penetrated into the target network and got to administrative control, and independently chose the further path on the found infrastructure, and did not perform a pre-prescribed scenario. In a more complex version, without credentials, Mythos also penetrated from the outside and reached the complete compromise of the domain. The final stage included obtaining the rights of Domain Admin and the ability to perform DCSync, that is, request domain secrets through the Active Directory replication mechanism.
Three other models, Grok-4.5, Muse Spark 1.1 and GLM-5.2, in separate launches also achieved full control over the domain, but could not pass the entire test sequence as steadily. GPT-5.6 Sol, Kimi K3, GPT-5.5-Cyber and DeepSeek-V4-Pro have reached the way between systems within the network. Almost all participants were able to at least gain initial access. The exception was Qwen3-Coder.
The most serious boundary was found not when using already known techniques, but when searching for new holes. In a simple test with deliberately implemented vulnerability, American, Chinese, open and closed models showed close results. Quite differently, the experiment with a real previously unknown error in a large fragment of the compiled software ended. Only Claude Mythos was able to understand the vulnerability and turn the find into a working operation. Such zero-day is especially valuable for the attacker, as the ready description and pre-written exploit at the model is not.
The study shows the limitation of the rating itself. With the main testing, Booz Allen intentionally did not give the models additional specialized "strapping" to compare the abilities directly of AI. A separate experiment showed how much the result changes the attack harness, that is, the software layer that connects the model to tools, memory, feedback and the logic of the attack. With such a system, even Claude Sonnet was able to get closer to Mythos. The authors believe that the real danger should therefore be assessed on the whole bundle of the model, tools and level of autonomy, rather than by the name of the model in the rating.
The index also does not prove that the listed models are already massively used for autonomous attacks on the Internet. Booz Allen worked in a controlled environment, and the indicators reflect the technical ability to perform separate offensive tasks. However, the shift from answering questions about security to self-selecting teams, finding the way through the network, and recovering from failed actions reduces the amount of human work needed for complex hacking.
Anthropic itself restricts access to the most powerful version of the line. Claude Mythos 5.1 is only available to proven organizations as part of special programs for cybersecurity and biology specialists. The company explicitly recognizes the dual purpose of the model and maintains a stricter access mode for it. Booz Allen considers the current gap temporary. If the researchers' prognosis is justified, the main indicator of the danger of AI for IS will no longer be the ability to write an exploit according to the instructions, but the ability without a person to link dozens of individual actions into a finished operation.
The gap between the first and second place was huge. Claude Mythos received 80 points, Grok-4.5 scored 49, GPT-5.6 Sol 46, Muse Spark 1.1 and Chinese Kimi K3 at 38, and GLM-5.2 from Z.ai 37. At the bottom of the ranking were Claude Sonnet 5 with 13 points, GLM-4.5-Air with 11, Qwen3.6-35B with 9 and Qwen3-Coder with 4. Booz Allen believes the backlog will shorten rapidly and most of the models tested will be able to approach Mythos capabilities within about six months. The latest forecast reflects the assessment of researchers, rather than the confirmed pace of model development.
The name Cyber Weapon Index sounds loud, but the methodology evaluates quite specific skills. The final score is equal to two indicators. Vulnerability Research Score shows how well the model self-founds vulnerabilities in compiled code without source. Kill Chain Attainment Score measures how far the model will advance along the attack chain in the secure enterprise Active Directory network. The formula is extremely simple, CWI is equal to the average value of two estimates.
For the second part of the test, the researchers gave each model a separate attacking machine and allowed to independently execute commands against the working test corporate network. The models were explored, gained initial access, fixed, mined credentials, increased privileges, performed lateral movement and tried to gain full control over the domain. The tests were carried out both with and without pre-provided stolen credentials. The result was verified by network traffic, host logs, domain controllers, and intrusion detection systems, so models could not get points for a simple statement about successful hacking.
Fraudsters learn every day. You are not.
SecurityLab in Telegram.
Claude Mythos has noticeably stood out on the practical part. With the current credentials, the model in all attempts penetrated into the target network and got to administrative control, and independently chose the further path on the found infrastructure, and did not perform a pre-prescribed scenario. In a more complex version, without credentials, Mythos also penetrated from the outside and reached the complete compromise of the domain. The final stage included obtaining the rights of Domain Admin and the ability to perform DCSync, that is, request domain secrets through the Active Directory replication mechanism.
Three other models, Grok-4.5, Muse Spark 1.1 and GLM-5.2, in separate launches also achieved full control over the domain, but could not pass the entire test sequence as steadily. GPT-5.6 Sol, Kimi K3, GPT-5.5-Cyber and DeepSeek-V4-Pro have reached the way between systems within the network. Almost all participants were able to at least gain initial access. The exception was Qwen3-Coder.
The most serious boundary was found not when using already known techniques, but when searching for new holes. In a simple test with deliberately implemented vulnerability, American, Chinese, open and closed models showed close results. Quite differently, the experiment with a real previously unknown error in a large fragment of the compiled software ended. Only Claude Mythos was able to understand the vulnerability and turn the find into a working operation. Such zero-day is especially valuable for the attacker, as the ready description and pre-written exploit at the model is not.
The study shows the limitation of the rating itself. With the main testing, Booz Allen intentionally did not give the models additional specialized "strapping" to compare the abilities directly of AI. A separate experiment showed how much the result changes the attack harness, that is, the software layer that connects the model to tools, memory, feedback and the logic of the attack. With such a system, even Claude Sonnet was able to get closer to Mythos. The authors believe that the real danger should therefore be assessed on the whole bundle of the model, tools and level of autonomy, rather than by the name of the model in the rating.
The index also does not prove that the listed models are already massively used for autonomous attacks on the Internet. Booz Allen worked in a controlled environment, and the indicators reflect the technical ability to perform separate offensive tasks. However, the shift from answering questions about security to self-selecting teams, finding the way through the network, and recovering from failed actions reduces the amount of human work needed for complex hacking.
Anthropic itself restricts access to the most powerful version of the line. Claude Mythos 5.1 is only available to proven organizations as part of special programs for cybersecurity and biology specialists. The company explicitly recognizes the dual purpose of the model and maintains a stricter access mode for it. Booz Allen considers the current gap temporary. If the researchers' prognosis is justified, the main indicator of the danger of AI for IS will no longer be the ability to write an exploit according to the instructions, but the ability without a person to link dozens of individual actions into a finished operation.