Skip to main content

Cyber AI Evaluation and Training

XOR is a cyber frontier lab. We build runnable security environments with executable verifiers, measure frontier models against them in public, and train on what the models do inside them.

See the results

Standards work by our researchers

Each agent gets a repository at the commit before a real fix landed, and has to repair the weakness. Nothing is graded by another model: a verifier replays the input that triggered the weakness and checks the code still builds.

Rates on this board are conditioned on frontier failure and are not comparable to a randomly drawn set.

1,581 scored attempts on the expert board · run 2026-05-08 · 3 boards, never pooled

Expert board

What is measured
An agent gets one repository at the commit before a fix, and the weakness it must resolve. A repair counts only when the verifier confirms the weakness no longer triggers, and the build still succeeds.
On what
100 environments across 16 agents, 1,581 scored attempts. Intervals resample codebases, not samples, because two weaknesses in one codebase are not independent draws.
How it was chosen
every frontier agent already failed the sample under both its own harness and a second oneRates on this board are conditioned on frontier failure and are not comparable to a randomly drawn set.

86 of 100 environments still tell two agents apart. 14 were repaired by no agent at all, and none were repaired by every agent. This board has no ceiling yet. An environment that everyone repairs, or that nobody repairs, separates nothing; the count between them is the board’s real size.

Each environment was attempted once here, so this board cannot say how much of a score would survive a rerun. That is unmeasured, not stable.

Loading the board…

Go deeper

Per-environment outcomes, cost per verified repair, rerun spread, and the full method live in the explorer. The data behind this page was regenerated 2026-08-26; the newest evaluation run in it is dated 2026-05-18.

Explore the data