Skip to main content
XOR

Benchmarks

Where we evaluate, and why

What we are testing

Inference runs on hardware a lab must trust. That trust rests on an attestation chain, the proof a machine gives of what it is running. A vulnerability in that chain is a vulnerability in the trust anchor.

Before publication

XOR handles discovery, disclosure and remediation end to end. At the time of these runs, no public fix or write-up was available.

Why we can be precise

XOR built the private confidential inference stack these runs use. Each run stays isolated on it, which is why we can say exactly what a patch had to hold against.

UNPATCHED · Attack leaderboard

Unpatched

Vulnerabilities XOR researchers discovered and disclosed in a confidential inference stack.

Ordered by Full exploit, highest first. Ties sort by Shallow exploit, then by name. Unranked.
Claude Code Opus 5 max95.0%0.0%
OpenCode Kimi K3 xhigh85.0%0.0%
OpenCode Qwen 3.8 Max xhigh84.2%0.0%
OpenCode GLM-5.3 xhigh78.9%0.0%
Codex GPT-5.6 Sol xhigh40.0%0.0%

In order of pass rate; too few runs to put one above another.

Unpatched leaderboardTop 5 of 6 systems measured
CRYPTOGRAPHY · Patch that keeps users in leaderboard

Cryptography

Planted vulnerabilities in cryptography libraries with no public fix: a patch must stop every attack and keep users in.

Ordered by Full fix, highest first. Ties sort by Shallow fix, then by name. Unranked.
OpenCode Gemini 3.8 Flash xhigh100.0%100.0%
OpenCode Kimi K3 xhigh100.0%100.0%
Claude Code Opus 5 max100.0%66.7%
OpenCode Qwen 3.8 Max xhigh100.0%66.7%
OpenCode GLM-5.3 xhigh66.7%66.7%

In order of pass rate; too few runs to put one above another.

Cryptography leaderboardTop 5 of 6 systems measured
ATTACK VERSUS DEFEND

Attack versus defend

Memory-safety vulnerabilities in low-level software with public fixes: defend rate minus attack rate.

← attackerdefender →
  1. OpenCode GPT-5.3 Codex: +6.7 points, -0.3 to +13.8, not separated
  2. OpenCode GPT-5.5: -11.0 points, -16.5 to -6.2, better as the attacker
  3. Claude Code Opus 4.7: -22.5 points, -30.8 to -13.9, better as the attacker
Defender rate minus attacker rate, in percentage points, on the vulnerabilities every system saw. Thin line: 95% CI. † The 95% CI includes zero.
AUTOMATED VULNERABILITY REPAIR

Vulnerability repair

Memory-safety vulnerabilities in low-level software with public fixes: systems patch them and exploit them.

On these leaderboards, exploit measures attacker performance while patch measures defender performance.

RankSystemPass rate
1
Codex GPT-5.5
86.3%80.5–90.6
1
OpenCode GPT-5.5
85.3%78.3–90.8
1
OpenCode GPT-5.3 Codex
84.0%78.0–88.6

Systems whose ranges overlap share a rank.

Vulnerability repair leaderboardTop 3 of 18 systems measured

Compare systems

Choose any measured leaderboard to compare up to five systems side by side.

Compare systems