Patch leaderboard · runs from April 2026 · as of 30 April 2026 · 18 systems
AUTOMATED VULNERABILITY REPAIR
Vulnerability repair
Each system repairs the vulnerability; the grader replays the trigger and rebuilds.
Exploit is the attacker role; Patch is the defender role.
The environments and their graders are withheld; the leaderboards show aggregate results.
Leaderboard
Leaderboard18 systems
Each bar is the share of vulnerabilities where the system's patch passed the grader.
- 1Codex GPT-5.586.3% 80.5–90.6, 95% CI 80.5–90.6
- 1OpenCode GPT-5.585.3% 78.3–90.8, 95% CI 78.3–90.8
- 1OpenCode GPT-5.3 Codex84.0% 78.0–88.6, 95% CI 78.0–88.6
- 1OpenCode Claude Opus 4.781.6% 73.9–87.4, 95% CI 73.9–87.4
- 1Codex GPT-5.481.3% 74.6–86.5, 95% CI 74.6–86.5
- 1OpenCode Claude Opus 4.679.6% 71.9–86.0, 95% CI 71.9–86.0
- 1OpenCode GPT-5.479.1% 70.5–85.2, 95% CI 70.5–85.2
- 1Gemini CLI 3.1 Pro Preview77.4% 69.9–83.5, 95% CI 69.9–83.5
- 1Cursor GPT-5.476.1% 67.2–83.1, 95% CI 67.2–83.1
- 1Cursor Claude Opus 4.675.7% 68.9–81.2, 95% CI 68.9–81.2
- 1Codex GPT-5.375.5% 67.1–81.6, 95% CI 67.1–81.6
- 1OpenCode GLM-5.175.1% 67.5–81.0, 95% CI 67.5–81.0
- 2Kiro Claude Opus 4.673.4% 65.7–79.2, 95% CI 65.7–79.2
- 4Cursor GPT-5.3 Codex71.3% 63.1–77.0, 95% CI 63.1–77.0
- 4OpenCode Kimi K2.670.2% 60.4–77.4, 95% CI 60.4–77.4
- 8Cursor Composer 263.2% 54.2–70.3, 95% CI 54.2–70.3
- 14Claude Code Opus 4.756.9% 48.1–65.2, 95% CI 48.1–65.2
- 16Claude Code Opus 4.646.2% 37.4–54.5, 95% CI 37.4–54.5
Leaderboard: Patch · newest run 30 April 2026 · 18 systems · Vulnerabilities were not picked for difficulty or by earlier system results. The reasoning effort for these runs is not recorded.
How we measured this
- Rank. One plus the number of systems whose whole range sits above this system's. Systems whose ranges overlap share a rank. A medal marks a rank held alone.
- Score. Share of complete runs where the patch passed the grader.
- Range. The thin line beside a score: the range the true rate probably sits in (its 95% confidence interval, CI). The unit of uncertainty is the codebase, because vulnerabilities in the same codebase are not independent draws.
- Coverage. OpenCode Claude Opus 4.7 ran a subset of the vulnerabilities.
- Harness versions. A system is a harness at a version with a model.
- Codex GPT-5.5
- 0.118.0
- OpenCode GPT-5.5
- 1.14.20
- OpenCode Claude Opus 4.7
- 1.14.20
- OpenCode Claude Opus 4.6
- 1.14.20
- Cursor GPT-5.4
- 2026.04.17-787b533
- Cursor Claude Opus 4.6
- 2026.04.17-787b533
- OpenCode GLM-5.1
- 1.14.20
- Kiro Claude Opus 4.6
- 2.0.1
- Cursor GPT-5.3 Codex
- 2026.04.17-787b533
- OpenCode Kimi K2.6
- 1.14.20
- Claude Code Opus 4.7
- 2.1.111
- Newest run 30 April 2026 · Cite · Method. Runs from April 2026, published 6 September 2026.