Benchmarks
Evaluation frameworks and datasets
Security and risk benchmarks 30
- PHARE by Giskardphare.giskard.ai
- RealHarm by Giskardrealharm.giskard.ai
- Inside the Benchmark: App Architectures, Finding Walkthroughs, and What Each Scanner Actually Caughtprojectdiscovery.io
- N-Day-Benchndaybench.winfunc.com
- AI Bug-Bounty Agent Benchmark: 10 Models, 100 Black-Box Labsmdpsec.com
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Studyarxiv.org
- Benchmarking 13 AI Models on Known CVE Detectionaikido.dev
- bots-bench - Benchmarking AI Models and Agents for SOC and IR Investigationsbotsbench.com
- BoxPwnr - benchmarking LLMs and agents on security CTF challengesgithub.com
- Cisco LLM Security Leaderboardleaderboard.aidefense.cisco.com
- Cotool BlueBench Windows Enterprise Intrusion benchmarkcotool.ai
- CTI-REALM: A new benchmark for end-to-end detection rule generation with AI agentsmicrosoft.com
- CVE-Bench - Benchmarking LLMs on real-world CVE patchinggiovannigatti.github.io
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacksarxiv.org
- FusionBench - Benchmark of AI Security Scan Models on a Fixed Vulnerability Corpusfusionbench.vercel.app
- Introducing dfbench v1: A cybersecurity benchmark for frontier models and agentic systemsdepthfirst.com
- Introducing EchoBench: A Human-Calibrated Benchmark for Autonomous Pentesting - NetSPInetspi.com
- JailbreakBench - LLM jailbreak robustness benchmarkjailbreakbench.github.io
- LLMs Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluationarxiv.org
- PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilitiesarxiv.org
- PWNBench: AI Pentesting Benchmark for 11 Frontier LLMsnovee.security
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potentialarxiv.org
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenariosarxiv.org
- SnitchBench - AI whistleblower benchmark visualizersnitchbench.t3.gg
- Snyk VulnBench - Benchmark of how reliably AI systems find vulnerabilitiesvulnbench.com
- Sol Searching - Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysissentinelone.com
- T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Modelsarxiv.org
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluationsarxiv.org
- We burned 11.7bn tokens to find the best cyber AI model - GLM5.3 and DeepSeek are now frontieraikido.dev
- We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks - Semgrept.co
Evaluation frameworks 20
- BattleBenchbattlebench.ai
- Microsoft Foundry risk and safety evaluations (preview) Transparenclearn.microsoft.com
- Minimum Viable Benchmarkblog.nilenso.com
- Benchmarking GPT-5.1 vs Gemini 3.0 v Opus 4.5 across 3 Coding Tasksblog.kilo.ai
- Alpha Arena - AI Trading Benchmarknof1.ai
- LLMs are beating regex in secret detection. We benchmarked Gitleaks, TruffleHog, andx.com
- LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Gooartificialanalysis.ai
- Long-form Content: AI (AI is Anti-Human (and assorted qualificationdocker.com
- Scheming reasoning evaluations — Apollo Researchapolloresearch.ai
- eyeballvul: a future-proof benchmark for vulnerability detection inarxiv.org
- Did You Train on My Dataset? Towards Public Dataset Protection witharxiv.org
- Arize Phoenix - open-source AI observability and evaluationgithub.com
- Dave Kennedy on Model Regression - daily benchmark testing of Claude, GPT, and Grokx.com
- Evaluating Large Language Models' Abilities to Process and Understand Technical Policy Reportsrand.org
- MatrAIx: Simulating the World with 8.3 Billion Persona Agentsarxiv.org
- ModelRegression.com - AI model performance and regression trackermodelregression.com
- Optimal stopping: spending evaluation compute where it countsaisi.gov.uk
- Run evaluations from the Microsoft Foundry portallearn.microsoft.com
- SAGE: A Generic Framework for LLM Safety Evaluationarxiv.org
- The Human Creativity Benchmarkcontralabs.com
Agent evaluation 10
- DoomArena: A framework for Testing AI Agents Against Evolving Security Threatsarxiv.org
- CyberGym: Evaluating AI Agents' Cybersecurity Capabilities with Real-World Vulnerabilities at Scalearxiv.org
- Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Designarxiv.org
- CAIBench - A Meta-Benchmark for Evaluating Cybersecurity AI Agentsarxiv.org
- CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysisarxiv.org
- Enclave: We Raced Seven AI Models to RCEenclave.ai
- Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skillsarxiv.org
- Mine the Gap: Open-Source Tools for Measuring the AI Offense-Defense Gapdreadnode.io
- SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agentsarxiv.org
- SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integrationarxiv.org