Cyber AI Evals
XOR is a security lab that evaluates and improves the cyber capabilities of coding systems.
See the resultsWhat we measure
A coding system meets a vulnerability with no public fix, and is set to attack it, repair it, or both. Code decides the outcome, and every result is published.
Benchmarks
Unpatched
0.0%Full exploit rateThe flaw fired at 63.5%; none cleared the bar for the full result.
Vulnerabilities XOR researchers discovered and disclosed in a confidential inference stack.
Unpatched benchmark →Cryptography
68.2%Full fix rate
Planted vulnerabilities in cryptography libraries with no public fix: a patch must stop every attack and keep users in.
Cryptography benchmark →[RESEARCH]
Cyber frontier eval: GPT-6 Astra and Claude Fable 5.1 repair only a third of the flaws in real-world conditions
Nine coding systems met flaws with no public fix. Two of them repaired about a third of what they were given, and no full exploit was accepted from any system on any run.
GPT-6 Astra stopped every attack by taking the service offline, and earned zero
Security fixes get rolled back for a dull reason more often than a dramatic one: they break the people who were supposed to keep working. The system under repair here is the broker between an engineer and production credentials, attacked in chained stages where each one opens the next. A repair has to stop every stage and still serve a legitimate caller. GPT-6 Astra shut the path the attacks used, shut its own operators out with it, and earned nothing.
We graded GLM-5.3-Flash before it had a name
In an earlier evaluation, four of seven systems repaired the vulnerability and then rejected every legitimate message, with every test still passing. The system that stopped every attack earned zero, because it stopped the legitimate traffic with it.
Method
Systems attack or defend, and code decides the outcome. Get in touch.