We ran Kimi K3 on a private cybersecurity benchmark.
TL;DR: Kimi K3 is the workhorse for cyber security tasks at great recall/precision/price. GPT 5.6 is best recall/precision but at 7x higher cost per run.
For context,
Deepsec.sh is an open-source cyber harness designed for finding vulnerabilities in large codebases.
The eval runs deepsec on an undisclosed open-core application at a git sha before a large number of security issues were fixed. This is a secret eval that cannot be directly benchmark-maxxed.
S-Tier: GPT 5.6 Sol: By far the most thorough analysis, but coming in at over 7x the price of the runner up.
Best price/recall: Kimi K3. Next tier of recall at a good price
Best price at good recall: GLM 5.2 (40% lower price than Kimi K3)
GPT 5.5: Only recommended with subscription or high-discount API price. Similar recall to Kimi at much higher list price.
Opus 4.8: Only recommended with subscription or high-discount API price. Similar recall to GLM 5.2 at much higher list price.
Fable 5: 100% refusal rate. Cannot be used for security analysis.
Sol on a large code base will quickly get into 6-figure pricing. This is still affordable relative to the risk of letting security issues unfixed or paying bug bounties.
I'd recommend using Sol for a one-time baseline and then using Kimi K3 for continuous analysis.
When using open-weight models, make sure to use an inference vendor that supports zero data retention.