Description
CVE-Bench Leaderboard is a platform designed to assess the capability of AI agents in autonomously exploiting web vulnerabilities. The evaluation is based on a dataset of 40 critical Common Vulnerabilities and Exposures (CVEs) announced by NIST between May 1, 2024, and June 14, 2024. These CVEs cover a wide range of high-stakes exploits, including remote code execution, SQL injection, and privilege escalation.
The platform provides two realistic settings for evaluation: the one-day setting, where vulnerability descriptions are provided, and the zero-day setting, where descriptions are not available. This allows for a comprehensive assessment of an AI agent's ability to handle both known and unknown vulnerabilities.
The leaderboard features various agents, such as the Default Agent combined with Claude Opus 4.6, achieving a Pass@1 rate of 32.5% at an average cost of $0.31 per task. Another example is the T-Agent combined with GPT-4o, which has a Pass@1 rate of 8.0% with an average cost of $1.70 per task.
CVE-Bench is particularly useful for organizations and researchers interested in understanding the effectiveness of AI in cybersecurity. It provides insights into the strengths and weaknesses of different AI agents in handling web vulnerabilities, making it a valuable tool for improving AI-driven security solutions.
CVE-Bench Leaderboard's Core Features
Evaluates AI agents on web vulnerabilities
Dataset of 40 critical CVEs
Covers remote code execution, SQL injection
One-day evaluation setting
Zero-day evaluation setting
Pass@1 performance metric
Cost per task analysis
Benchmark version tracking
How to use CVE-Bench Leaderboard?
Select evaluation setting: choose one-day or zero-day
Run AI agent: execute the agent on selected CVEs
Analyze results: review Pass@1 and cost metrics
Optimize agent: improve based on performance data
CVE-Bench Leaderboard's Use Cases
- AI vulnerability assessment
- Cybersecurity research
- AI agent optimization
- Security solution development
- Benchmarking AI tools








