Skip to main content

gpt-5.2-codex vs opus-4.6

Both models ran the same corpus under the same verifier — codex and cursor respectively. The harness differs where the names differ, and that is part of what separates them.

Their confidence intervals overlap. On this corpus the difference of 13.3 points is not resolved — treat these two as tied.

codex / gpt-5.2-codex

49.2%

95% interval 35.9–59.8 · 128 attempts

49.2% (95% interval 35.9–59.8)
Full profile →

cursor / opus-4.6

62.5%

95% interval 50.7–71.9 · 128 attempts

62.5% (95% interval 50.7–71.9)
Full profile →

Where they differ

Weaknessgpt-5.2-codexopus-4.6Difference
CWE-125Out-of-bounds Read56.0%65.3%-9.3
CWE-787Out-of-bounds Write64.7%76.5%-11.8
CWE-416Use After Free15.4%46.2%-30.8
CWE-908Use of Uninitialized Resource22.2%33.3%-11.1
CWE-415Double Free33.3%66.7%-33.4
CWE-1284Improper Validation of Specified Quantity in Input50.0%50.0%0.0
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0%50.0%0.0
CWE-476NULL Pointer Dereference0.0%50.0%-50.0
CWE-475100.0%100.0%0.0
CWE-5620.0%100.0%-100.0

Per-class samples are small; a difference here is a hint about where to look, not a result. The corpus-level intervals above are the honest comparison.

Corpus cve-bench-136 · run February 2026 · How we test

See the benchmark →