SenseNova (China)
Aikido's cyber-capability benchmark burned 11.7 billion tokens across 10 AI models, each running 32 fresh off-the-shelf vulnerabilities three times. DeepSeek V4 Pro 0813 was the top performer, pooling three runs to find 28 of 32 vulnerabilities, with three DeepSeek Pro runs costing approximately $295 and outperforming The benchmark reports that single runs miss much of the breadth, but pooling across runs fills that gap: DeepSeek Pro finds 17 vulnerabilities on its first pass but 28 across three. It concludes that open-source models now outperform the public frontier on pooled vulnerability recall, with DeepSeek V4 Pro topping every