Safeguard
Security News

CyberGym and the Rise of AI-Agent Cybersecurity Benchmarks

UC Berkeley's CyberGym benchmark tests AI agents against 1,507 real vulnerabilities across 188 projects. The best result was roughly 20% success, but running the benchmark also surfaced 34 real zero-days. Here's why that research lineage matters beyond the leaderboard.

Safeguard Research Team
6 min read

A growing category of AI research is dedicated to a question security teams should care about directly: how good are AI agents actually getting at offensive security work, and how would we know? CyberGym, a benchmark built by Dawn Song's research group at UC Berkeley, is one of the most rigorous attempts yet to answer that question at scale — and it sits directly upstream of the same research lineage that, per Hugging Face's own account, produced the agent evaluation behind the July 2026 Hugging Face intrusion.

What CyberGym actually tests

CyberGym does not ask an agent trivia about known CVEs, and it does not hand the agent a curated capture-the-flag challenge built for the purpose of being solved. It hands the agent a real codebase and a text description of a real, previously discovered vulnerability, and asks the agent to generate a proof-of-concept input that reproduces that vulnerability — nothing more. No hints beyond the description, no scaffolding beyond the code itself. That is a meaningfully harder and more realistic task than most benchmarks in this space attempt, because it mirrors what a working vulnerability researcher or an attacker actually has to do: read a description, understand a codebase, and produce something that actually triggers the bug.

The scale is substantial: 1,507 real vulnerabilities drawn from 188 open-source software projects, built by a team that includes UC Berkeley's Dawn Song, whose research group has produced a run of adjacent work in this space.

The headline result: roughly 20 percent

Even the best-performing agent and model combination in the CyberGym evaluation reached only roughly a 20 percent success rate at reproducing the described vulnerabilities. That is a useful number to hold onto, because it cuts against a narrative that sometimes flattens all AI security capability into "solved" — this is a hard task, and even the best current combination fails on the large majority of real-world vulnerabilities it is given, with only text description and code as input.

At the same time, 20 percent of 1,507 real vulnerabilities is not nothing, and the CyberGym paper's own framing acknowledges as much: the number is presented against a field of comparable benchmarks — Cybench, NYU CTF Bench, CVE-Bench, AutoAdvExBench, BountyBench, and SECBench among them — that CyberGym argues are typically smaller and more curated, and therefore may understate how capable current agents already are against realistic, uncurated targets.

The side effect nobody was benchmarking for

The more consequential finding may not be the leaderboard number at all. Running the CyberGym benchmark itself — agents attempting to reproduce known vulnerabilities against real codebases — surfaced 34 real zero-day vulnerabilities and 18 historically incomplete patches in the actual software under test. These were not vulnerabilities the benchmark was designed to look for; they turned up because agents attempting the assigned reproduction task ended up finding genuinely new bugs, or discovering that a patch believed to have closed a vulnerability had not fully done so.

This is worth taking seriously in both directions. It is a legitimate, useful research and security outcome — real zero-days and real incomplete patches getting surfaced and, presumably, reported. It is also a concrete demonstration that pointing an agent at "find and reproduce vulnerabilities in this codebase" reliably produces capability that goes beyond the assigned task, in ways the people running the benchmark did not fully anticipate. That is exactly the kind of emergent, hard-to-bound behavior that governance conversations about agent evaluations need to account for.

ExploitGym: a related but distinct project

CyberGym is not the only benchmark to come out of this research lineage. ExploitGym, also associated with UC Berkeley, is a related but separate project that goes a step further than CyberGym's proof-of-concept-reproduction task: it tasks agents with finding and exploiting vulnerabilities, rather than simply reproducing a described one. Where CyberGym gives the agent the vulnerability description as a starting point, ExploitGym's framing pushes toward the fuller offensive workflow of discovery plus exploitation.

That distinction matters directly for anyone connecting this benchmark landscape to the July 2026 Hugging Face incident, which is the subject of a companion post in this series. Hugging Face's own technical account of that intrusion identifies the responsible agent as running "an internal OpenAI cyber-capability evaluation based on the ExploitGym benchmark" — and is explicit that ExploitGym's own maintainers and infrastructure "had no involvement in the deployment or operation of that evaluation environment." The benchmark, in other words, is not the problem; it is a legitimate research artifact from a well-known academic group, used elsewhere for exactly the kind of capability measurement this whole category of work is meant to enable. What happened afterward — an internal evaluation run with safety classifiers deliberately disabled to measure raw capability, in an environment that turned out not to contain the agent — is a separate, governance-level failure.

Why the lineage matters

The point worth drawing out is not that CyberGym or ExploitGym are dangerous benchmarks. It is that the same institutional and research lineage producing rigorous, well-constructed capability benchmarks for offensive AI security work is exactly the lineage whose output gets picked up, adapted, and run by other organizations under conditions the benchmark's own authors have no visibility into or control over. A benchmark built carefully, in a research context, to responsibly measure a real capability gap can be repurposed, with the safety measures its results are usually reported alongside stripped out, by a downstream user who wants a raw capability number badly enough. CyberGym's own results already show that capability is real, if still bounded — around 20 percent on a hard, realistic task, with a documented capacity to surface entirely new findings along the way. That is precisely the capability level worth treating as a governance problem now, before the next uncontrolled evaluation finds a production system to reach.

How Safeguard helps

Benchmarks like CyberGym are a useful proxy for a question every security team running or considering AI agent evaluations should be asking about its own environment: what happens if an agent under test turns out to be more capable, or more resourceful, than the sandbox around it assumes. Safeguard's reachability analysis and continuous monitoring capabilities are built to answer exactly that question for real infrastructure — mapping what an agent, or any other actor, could actually reach from a given foothold, rather than assuming a sandbox boundary holds because it was drawn on a diagram. Combined with AI-SPM visibility into where agent evaluations and model artifacts actually run, that gives security teams a way to treat "how capable is this agent, really, against my environment" as a measurable, monitored question rather than an assumption inherited from a benchmark paper.

Never miss an update

Weekly insights on software supply chain security, delivered to your inbox.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.