Which coding agent fixes real vulnerabilities?
XOR runs AI coding agents against 128 real memory-safety vulnerabilities in production systems code, and verifies every fix against the weakness it claims to resolve. 1,864 scored attempts, last evaluated February 2026 — 6 months ago. There is no fixed refresh cadence.
Agent is the tool the model runs inside; model is the model itself. The same model scores differently in different tools — by up to 14 points. Rows sharing a rank are tied: their confidence intervals overlap, so this data does not separate them.
| Rank | Agent | Model | Pass rate | 95% interval | Attempts |
|---|---|---|---|---|---|
| 1 | codex | gpt-5.2 | 62.7% | 51.9–71.5 | 126 |
| cursor | opus-4.6 | 62.5% | 50.7–71.9 | 128 | |
| claude | claude-opus-4-6 | 61.6% | 48.9–71.3 | 125 | |
| gemini31 | gemini-3.1-pro-preview | 58.7% | 44.0–71.2 | 109 | |
| opencode | gemini-3.1-pro-preview | 54.9% | 42.9–66.7 | 122 | |
| cursor | gpt-5.2 | 51.6% | 38.1–63.6 | 122 | |
| opencode | gpt-5.2 | 51.6% | 39.5–63.1 | 122 | |
| cursor | gpt-5.3-codex | 50.4% | 36.2–61.2 | 127 | |
| codex | gpt-5.2-codex | 49.2% | 35.9–59.8 | 128 | |
| opencode | claude-opus-4-6 | 47.5% | 32.0–59.7 | 122 | |
| claude | claude-opus-4-5 | 45.7% | 33.1–55.5 | 127 | |
| cursor | composer-1.5 | 45.2% | 32.5–58.0 | 126 | |
| gemini | gemini-3-pro-preview | 43.0% | 32.8–52.1 | 128 | |
| 2 | opencode | gpt-5.2-codex | 37.8% | 25.0–48.7 | 127 |
| opencode | claude-opus-4-5 | 36.8% | 25.0–48.5 | 125 |
This corpus separates 15 configurations into only 2 bands: within a band the confidence intervals overlap and the order above carries no information. That is the honest resolution of this benchmark, not a defect in it.
128 samples · 15 agent configurations · attempts vary by row (109–128) · How we test
Trajectories, per-sample outcomes, cost per verified fix, and the full methodology live in the benchmark explorer.
Explore the benchmark