Attack versus defend · runs from April 2026 · published 6 September 2026 · 3 systems
ATTACK VERSUS DEFEND
Attack versus defend
Memory-safety vulnerabilities in low-level software with public fixes: defend rate minus attack rate.
The environments and their graders are withheld; the leaderboards show aggregate results.
Paired outcomes
3 systems · the vulnerabilities every one of them finished in both rolesA paired view of the April 2026 vulnerability repair data, so the systems here differ from the ones on the September benchmarks.
← better as the attackerbetter as the defender →
- OpenCode GPT-5.3 Codex: +6.7 points, -0.3 to +13.8, not separated
- OpenCode GPT-5.5: -11.0 points, -16.5 to -6.2, better as the attacker
- Claude Code Opus 4.7: -22.5 points, -30.8 to -13.9, better as the attacker
The table scrolls sideways.
| System | Δ (pp) | Range | Attack | Defend |
|---|---|---|---|---|
| OpenCode GPT-5.3 Codex | +6.7 (the range includes zero) | -0.3 to +13.8 | 77.0% | 83.7% |
| OpenCode GPT-5.5 | -11.0 | -16.5 to -6.2 | 96.2% | 85.2% |
| Claude Code Opus 4.7 | -22.5 | -30.8 to -13.9 | 79.9% | 57.4% |
Paired set · newest run 30 April 2026 † the range includes zero. The reasoning effort for these runs is not recorded. On one harness, GPT-5.3 Codex minus GPT-5.5: +17.7 pp, +10.5 to +24.9. Claude Code ran CLI 2.1.114 as attacker and 2.1.111 as defender, so this differential does not attribute to the model alone. Paired across a small harness version difference; by our definition these are close but not identical systems.