Cyber AI Evaluation
XOR is a security lab that measures coding systems as attacker and as defender on vulnerabilities with no public fix, and publishes what it finds.
See the resultsWhat we measure
We give a coding system a vulnerability nobody has fixed yet, and ask it to attack, repair, or both. Code decides the outcome, and every result goes on a public leaderboard.
Compare system resultsBenchmarks
Unpatched
0.0%Graded runs with a full exploit
Vulnerabilities XOR researchers discovered and disclosed in confidential-computing attestation libraries within a confidential inference stack.
Unpatched benchmark →Cryptography
68.2%Graded runs with a full fix
Planted vulnerabilities in cryptography libraries with no public fix: a patch must stop every attack and keep users in.
Cryptography benchmark →[RESEARCH]
We compared Unpatched outcomes
Attack and repair success on Unpatched: grading criteria matter.
The fix that locked everyone out
When one system stopped every planted attack by locking down the service entirely, our legitimate-use check caught the lockout and the grader awarded zero.
We graded GLM-5.3-Flash before it had a name
Our earlier evaluation found GLM-5.3-Flash preserving legitimate traffic while other patches blocked it despite passing tests.
Method
Systems attack or defend. How we measure.