Skip to main content
XOR
ATTACK VERSUS DEFEND

Attack versus defend

Memory-safety vulnerabilities in low-level software with public fixes: defend rate minus attack rate.

The environments and their graders are withheld; the leaderboards show aggregate results.

Paired outcomes

3 systems · the vulnerabilities every one of them finished in both roles

A paired view of the April 2026 vulnerability repair data, so the systems here differ from the ones on the September benchmarks.

attackerdefender →
  1. OpenCode GPT-5.3 Codex: +6.7 points, -0.3 to +13.8, not separated
  2. OpenCode GPT-5.5: -11.0 points, -16.5 to -6.2, better as the attacker
  3. Claude Code Opus 4.7: -22.5 points, -30.8 to -13.9, better as the attacker
Defender rate minus attacker rate, in percentage points, on the vulnerabilities every system saw. Thin line: 95% CI. † The 95% CI includes zero.

The table scrolls sideways.

Per system: the defend rate minus the attack rate, its range, † where that range includes zero, then the attack and defend rates.
SystemΔ (pp)RangeAttackDefend
OpenCode GPT-5.3 Codex+6.7 (the range includes zero)-0.3 to +13.877.0%83.7%
OpenCode GPT-5.5-11.0-16.5 to -6.296.2%85.2%
Claude Code Opus 4.7-22.5-30.8 to -13.979.9%57.4%
Paired set · newest run 30 April 2026 † the range includes zero. The reasoning effort for these runs is not recorded. On one harness, GPT-5.3 Codex minus GPT-5.5: +17.7 pp, +10.5 to +24.9. Claude Code ran CLI 2.1.114 as attacker and 2.1.111 as defender, so this differential does not attribute to the model alone. Paired across a small harness version difference; by our definition these are close but not identical systems.
Attack versus defend | XOR