AI Security Capabilities Now Quantified, Too: ‘Cyber Index’ Debuts
IT DAILY ·
✦ AI Summary
Artificial Analysis unveiled a new metric, the “Cyber Index,” on the 28th, local time, to quantify AI cybersecurity capabilities. The index evaluates the full process of vulnerability discovery, validation, and remediation.
In the first evaluation, SpaceXAI’s “Grok 4.7” and Xiaomi’s “Mimo V2.6 Pro” tied for first place with 56 points. OpenAI’s “GPT-6 Luna” inference “Max” setting scored 53 points, and Z.ai’s “GLM 5.3 Flash” scored 50.
Some frontier AI models declined tasks or were blocked from testing because of safety guardrail issues, and Artificial Analysis assigned 0 points to models that refused. The index was co-developed by Artificial Analysis with IBM, NVIDIA, Collinear AI, and Vercel.
Artificial Analysis unveiled a new metric on the 28th, local time, to quantify AI’s cybersecurity capabilities: the “Cyber Index.” The “Cyber Index” is a benchmark for evaluating AI model capabilities in cybersecurity, and its scope covers the full process of vulnerability discovery, validation, and remediation.
In the first evaluation released by Artificial Analysis, SpaceXAI’s “Grok 4.7” and Xiaomi’s “Mimo V2.6 Pro” tied for first place. The shared No. 1 score was 56 points.
However, some frontier models from major AI companies repeatedly declined certain evaluation tasks for safety reasons. Frontier AI models such as GPT-6 Astra were blocked from testing because of safety guardrail issues, and those models could not be evaluated. This task refusal and test blocking among frontier models exposed the limits of capability comparisons.
Artificial Analysis co-developed the index with IBM, NVIDIA, Collinear AI, and Vercel. In the process, IBM and NVIDIA provided expert input on methodology, while Collinear AI and Vercel participated in the evaluation process. The collaboration also launched the “Cyber Index Alliance,” a consortium.
The Cyber Index is calculated by giving equal weight to CWE-Bench-AA, DeepSecBench-AA, and CyberGym-E2E-AA. Each benchmark measures different defensive capabilities, such as vulnerability detection, validation, and patch development, and the final results are presented on a 0-100 point scale. The evaluation targets vulnerability detection and remediation in source-code-accessible environments, and excludes incident response from its scope.
Artificial Analysis unveiled the “Cyber Index” on the 28th, local time. This evaluation compared model Cyber Index scores and cost per task, and the index measured security capabilities and token cost.
In terms of scores, SpaceXAI’s “Grok 4.7” with the xhigh setting and Xiaomi’s “Mimo V2.6 Pro” tied for first in the initial public evaluation. The two models each scored 56 on the Cyber Index. OpenAI’s “GPT-6 Luna” inference “Max” setting followed with 53 points, while Z.ai’s “GLM 5.3 Flash” scored 50.
On cost, OpenAI’s “GPT-6 Luna” inference “Max” setting posted the lowest cost per task at USD 0.12. Xiaomi’s “Mimo V2.6 Pro” tied for first overall while costing USD 0.18 per task. Even among top-ranked models, cost efficiency varied widely, and SpaceXAI’s “Grok 4.7,” despite being among the highest-scoring models, had the highest cost of all evaluated systems at USD 11.67 per task.
The Cyber Index was presented as a meaningful metric for verifying AI model security capabilities. However, frontier AI was not included in this announcement. That was because the safety guardrails of model and service providers blocked some vulnerability-validation tasks. OpenAI’s “GPT-6 Astra” refused about 98% of CyberGym-E2E-AA benchmark tasks, and Anthropic’s “Claude Fable 5.1” also refused about 98% of CyberGym-E2E-AA benchmark tasks. Alibaba’s “Qwen 3.8 2.4T A95B” and “Qwen 3.8 27B” blocked all benchmark tasks.
In response, Artificial Analysis assigned a score of 0 to models that refused the tasks. It also separately indicated refusal rates.
Artificial Analysis, founded in 2023, is a global benchmark firm that evaluates the full AI stack, including language models, agent intelligence, images, video, and voice. The organization analyzes more than 500 models, and its results are cited by major AI companies.
Source: IT DAILY · Kim Ho-jun
Original: https://www.itdaily.kr/news/articleView.html?idxno=241886
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.