Skip to main content

Which coding agent fixes real vulnerabilities?

XOR runs AI coding agents against 128 real memory-safety vulnerabilities in production systems code, and verifies every fix against the weakness it claims to resolve. 1,864 scored attempts, last evaluated February 2026 — 6 months ago. There is no fixed refresh cadence.

Agent is the tool the model runs inside; model is the model itself. The same model scores differently in different tools — by up to 14 points. Rows sharing a rank are tied: their confidence intervals overlap, so this data does not separate them.

RankAgentModelPass rate95% intervalAttempts
1codexgpt-5.262.7%51.9–71.5126
cursoropus-4.662.5%50.7–71.9128
claudeclaude-opus-4-661.6%48.9–71.3125
gemini31gemini-3.1-pro-preview58.7%44.0–71.2109
opencodegemini-3.1-pro-preview54.9%42.9–66.7122
cursorgpt-5.251.6%38.1–63.6122
opencodegpt-5.251.6%39.5–63.1122
cursorgpt-5.3-codex50.4%36.2–61.2127
codexgpt-5.2-codex49.2%35.9–59.8128
opencodeclaude-opus-4-647.5%32.0–59.7122
claudeclaude-opus-4-545.7%33.1–55.5127
cursorcomposer-1.545.2%32.5–58.0126
geminigemini-3-pro-preview43.0%32.8–52.1128
2opencodegpt-5.2-codex37.8%25.0–48.7127
opencodeclaude-opus-4-536.8%25.0–48.5125

This corpus separates 15 configurations into only 2 bands: within a band the confidence intervals overlap and the order above carries no information. That is the honest resolution of this benchmark, not a defect in it.

128 samples · 15 agent configurations · attempts vary by row (109128) · How we test

Trajectories, per-sample outcomes, cost per verified fix, and the full methodology live in the benchmark explorer.

Explore the benchmark