Table of Contents
ToggleNew tests show that leading US artificial intelligence models still hold a clear lead over China’s Kimi K3 in offensive cybersecurity skills.
A joint evaluation by the UK Artificial Intelligence Security Institute and the US Center for AI Standards and Innovation found a wide performance gap.
The results come after Kimi K3 gained attention for strong results on general AI benchmarks. These new findings focus specifically on cybersecurity capabilities.
Cybersecurity Benchmark Results

Researchers used ExploitBench, a public test created by Carnegie Mellon University. The benchmark measures how well an AI model can develop complete software exploits.
Kimi K3 scored 32.2% overall on the cyber capability test. Top US models averaged 76.2%. Another Chinese open-weight model, GLM-5.2, scored even lower at 24%.
The evaluation also checked for arbitrary code execution, a serious type of exploit that lets attackers control a system. Kimi K3 achieved this on none of the 41 test tasks. The strongest US models succeeded on an average of 20 out of 41 tasks.
In a separate simulated attack called “The Last Ones,” Kimi K3 reached step 17 of a 32-step network attack on average. Leading US models reached step 28.5. Kimi K3 fully completed the attack in only one out of ten attempts.
What the Results Mean
The findings show that Kimi K3 can still carry out some autonomous cyber operations against weak systems when given access and clear instructions. Its built-in safeguards did not fully stop it from attempting these tasks during testing.
However, the overall gap remains large. Experts say the results provide a more balanced view of the current AI competition. While China has made progress in general AI performance, US models continue to lead in specialized cybersecurity capabilities.
David Sacks, a key US AI adviser, commented on the results. He said the findings suggest that concerns about Chinese models overtaking American systems may be overstated. He urged policymakers not to overreact with heavy regulation.
The evaluation is still preliminary. Researchers noted that Kimi K3’s score came from a single main benchmark, while other models were tested across more tasks. Further testing will give a clearer picture over time.
For now, the results confirm that top US AI systems maintain a significant advantage in offensive cyber skills.
Quick Links: