Cyber AI Evals
XOR is a security lab that evaluates and improves the cyber capabilities of coding systems.
See the resultsWhat we measure
A coding system meets a vulnerability with no public fix, and is set to attack it, repair it, or both. Code decides the outcome, and every result is published.
Benchmarks
Unpatched
0.0%Full exploit rateThe flaw fired at 63.5%; none cleared the bar for the full result.
Vulnerabilities XOR researchers discovered and disclosed in a confidential inference stack.
Unpatched benchmark →Cryptography
68.2%Full fix rate
Planted vulnerabilities in cryptography libraries with no public fix: a patch must stop every attack and keep users in.
Cryptography benchmark →[RESEARCH]
Cyber frontier eval: GPT-6 Astra and Claude Fable 5.1 repair only a third of the flaws in real-world conditions
Nine coding systems met flaws with no public fix. Two of them repaired about a third of what they were given, and no full exploit was accepted from any system on any run.
GPT-6 Astra stopped every attack by taking the service offline, and earned zero
Security fixes get rolled back for a dull reason more often than a dramatic one: they break the people who were supposed to keep working. The system under repair here is the broker between an engineer and production credentials, attacked in chained stages where each one opens the next. A repair has to stop every stage and still serve a legitimate caller. GPT-6 Astra shut the path the attacks used, shut its own operators out with it, and earned nothing.
We graded GLM-5.3 before it had a name
In an earlier evaluation, five of nine systems repaired the vulnerability and then rejected every legitimate message, with every test still passing. The one system that kept traffic working was the one we graded before it had a public name.
Method
Systems attack or defend. How we measure.