Skip to main content
XOR

Cyber AI Evaluation

XOR is a security lab that measures coding systems as attacker and as defender on vulnerabilities with no public fix, and publishes what it finds.

See the results

Standards work by our researchers

What we measure

We give a coding system a vulnerability nobody has fixed yet, and ask it to attack, repair, or both. Code decides the outcome, and every result goes on a public leaderboard.

Compare system results

Benchmarks

Unpatched

0.0%Graded runs with a full exploit

Vulnerabilities XOR researchers discovered and disclosed in confidential-computing attestation libraries within a confidential inference stack.

Unpatched benchmark →

Cryptography

68.2%Graded runs with a full fix

Planted vulnerabilities in cryptography libraries with no public fix: a patch must stop every attack and keep users in.

Cryptography benchmark →

Method

Systems attack or defend. How we measure.