SEC-bench Pro

Can Language Models Solve Long-Horizon Software Security Tasks?

SEC-bench Pro is a self-evolving security benchmark that measures agents' ability to hunt security bugs in critical software systems. More projects and security tasks will be added over time. Use Version to switch between benchmark snapshots.

Linux evaluates kernel vulnerability discovery and PoC generation. This track includes 137 source-file instances.

# Model Success Completed Provider Backend
1
77.4%
106/137
125/137 12 timed out OpenAI OpenAI
2
52.6%
72/137
124/137 13 timed out OpenAI OpenAI
3
Opus 4.6 (max) Claude Code
39.4%
54/137
57/137 80 timed out Anthropic AWS Bedrock
4
GLM-5 (high) OpenCode
Open
3.6%
5/137
134/137 3 timed out Z.ai AWS Bedrock
5
Kimi K2.5 (high) OpenCode
Open
2.2%
3/137
137/137 0 timed out Moonshot AI AWS Bedrock
6
Open
1.5%
2/137
132/137 5 timed out MiniMax AWS Bedrock

Newer results for Claude Opus and Mythos models are not available due to safeguard restrictions.

News

  • [June 17, 2026]: Overall and Linux leaderboards are now built.
  • [May 5, 2026]: SpiderMonkey leaderboard has been released.
  • [May 1, 2026]: SEC-bench Pro launches with the V8 leaderboard.

Overview

SEC-bench Pro evaluates agents on vulnerability discovery and PoC generation tasks on challenging targets. Each target instance ships with a Docker image, metadata, a rendered prompt, and a harness. The harness renders the prompt from meta.json, starts the instance container, runs the configured agent with a timeout, and collects the generated files and run artifacts for checker evaluation.