Cyber AI Evaluation and Training
XOR is a cyber frontier lab. We build runnable security environments with executable verifiers, measure frontier models against them in public, and train on what the models do inside them.
See the resultsStandards work by our researchers
What we measure
Fixing a vulnerability is the last of three jobs. An agent has to find the weakness, reach it, then repair it without breaking anything. We measure reach and repair. We do not measure finding.
- 01Find
Locate a weakness nobody has described.
No board here asks an agent to discover an unknown weakness.
[not measured]
- 02Reach
Trigger a weakness that has already been described.
The agent is handed the weakness and asked to reach it. That is reproduction, not discovery, and the rate should not be read as evidence an agent could find it unaided.
72.2%
of attempts produced a working trigger
1,268 scored · 6 agents · 212 environments
- 03Repair
Fix the weakness without breaking the build.
Across our three boards the best agent scores between 45% and 85%. The spread is the selection rule, not the capability: the boards are drawn differently and are never pooled.
45.0%
best agent on the hardest board
3,888 scored · 16 agents · 100 environments
The board
Each agent gets a repository at the commit before a real fix landed, and has to repair the weakness. Nothing is graded by another model: a verifier replays the input that triggered the weakness and checks the code still builds.
Rates on this board are conditioned on frontier failure and are not comparable to a randomly drawn set.
1,581 scored attempts on the expert board · run 2026-05-08 · 3 boards, never pooled
Expert board
- What is measured
- An agent gets one repository at the commit before a fix, and the weakness it must resolve. A repair counts only when the verifier confirms the weakness no longer triggers, and the build still succeeds.
- On what
- 100 environments across 16 agents, 1,581 scored attempts. Intervals resample codebases, not samples, because two weaknesses in one codebase are not independent draws.
- How it was chosen
- every frontier agent already failed the sample under both its own harness and a second oneRates on this board are conditioned on frontier failure and are not comparable to a randomly drawn set.
86 of 100 environments still tell two agents apart. 14 were repaired by no agent at all, and none were repaired by every agent. This board has no ceiling yet. An environment that everyone repairs, or that nobody repairs, separates nothing; the count between them is the board’s real size.
Each environment was attempted once here, so this board cannot say how much of a score would survive a rerun. That is unmeasured, not stable.
Loading the board…
Go deeper
Per-environment outcomes, cost per verified repair, rerun spread, and the full method live in the explorer. The data behind this page was regenerated 2026-08-26; the newest evaluation run in it is dated 2026-05-18.
Explore the data