Skip to main content
XOR
AUTOMATED VULNERABILITY REPAIR

Vulnerability repair

Each system repairs the vulnerability; the grader replays the trigger and rebuilds.

Exploit is the attacker role; Patch is the defender role.

The environments and their graders are withheld; the leaderboards show aggregate results.

Leaderboard
Leaderboard18 systems

Each bar is the share of vulnerabilities where the system's patch passed the grader.

  • 1
    Codex GPT-5.5
    86.3% 80.5–90.6, 95% CI 80.5–90.6
  • 1
    OpenCode GPT-5.5
    85.3% 78.3–90.8, 95% CI 78.3–90.8
  • 1
    OpenCode GPT-5.3 Codex
    84.0% 78.0–88.6, 95% CI 78.0–88.6
  • 1
    OpenCode Claude Opus 4.7
    81.6% 73.9–87.4, 95% CI 73.9–87.4
  • 1
    Codex GPT-5.4
    81.3% 74.6–86.5, 95% CI 74.6–86.5
  • 1
    OpenCode Claude Opus 4.6
    79.6% 71.9–86.0, 95% CI 71.9–86.0
  • 1
    OpenCode GPT-5.4
    79.1% 70.5–85.2, 95% CI 70.5–85.2
  • 1
    Gemini CLI 3.1 Pro Preview
    77.4% 69.9–83.5, 95% CI 69.9–83.5
  • 1
    Cursor GPT-5.4
    76.1% 67.2–83.1, 95% CI 67.2–83.1
  • 1
    Cursor Claude Opus 4.6
    75.7% 68.9–81.2, 95% CI 68.9–81.2
  • 1
    Codex GPT-5.3
    75.5% 67.1–81.6, 95% CI 67.1–81.6
  • 1
    OpenCode GLM-5.1
    75.1% 67.5–81.0, 95% CI 67.5–81.0
  • 2
    Kiro Claude Opus 4.6
    73.4% 65.7–79.2, 95% CI 65.7–79.2
  • 4
    Cursor GPT-5.3 Codex
    71.3% 63.1–77.0, 95% CI 63.1–77.0
  • 4
    OpenCode Kimi K2.6
    70.2% 60.4–77.4, 95% CI 60.4–77.4
  • 8
    Cursor Composer 2
    63.2% 54.2–70.3, 95% CI 54.2–70.3
  • 14
    Claude Code Opus 4.7
    56.9% 48.1–65.2, 95% CI 48.1–65.2
  • 16
    Claude Code Opus 4.6
    46.2% 37.4–54.5, 95% CI 37.4–54.5
Leaderboard: Patch · newest run 30 April 2026 · 18 systems · Vulnerabilities were not picked for difficulty or by earlier system results. The reasoning effort for these runs is not recorded.
How we measured this
  • Rank. One plus the number of systems whose whole range sits above this system's. Systems whose ranges overlap share a rank. A medal marks a rank held alone.
  • Score. Share of complete runs where the patch passed the grader.
  • Range. The thin line beside a score: the range the true rate probably sits in (its 95% confidence interval, CI). The unit of uncertainty is the codebase, because vulnerabilities in the same codebase are not independent draws.
  • Coverage. OpenCode Claude Opus 4.7 ran a subset of the vulnerabilities.
  • Harness versions. A system is a harness at a version with a model.
    Codex GPT-5.5
    0.118.0
    OpenCode GPT-5.5
    1.14.20
    OpenCode Claude Opus 4.7
    1.14.20
    OpenCode Claude Opus 4.6
    1.14.20
    Cursor GPT-5.4
    2026.04.17-787b533
    Cursor Claude Opus 4.6
    2026.04.17-787b533
    OpenCode GLM-5.1
    1.14.20
    Kiro Claude Opus 4.6
    2.0.1
    Cursor GPT-5.3 Codex
    2026.04.17-787b533
    OpenCode Kimi K2.6
    1.14.20
    Claude Code Opus 4.7
    2.1.111
  • Newest run 30 April 2026 · Cite · Method. Runs from April 2026, published 6 September 2026.
Vulnerability repair | XOR