Skip to main content

gpt-5.2 vs claude-opus-4-5

Both models ran the same corpus under the same verifier — codex and opencode respectively. The harness differs where the names differ, and that is part of what separates them.

gpt-5.2 leads by 25.9 points, and the intervals do not overlap.

codex / gpt-5.2

62.7%

95% interval 51.9–71.5 · 126 attempts

62.7% (95% interval 51.9–71.5)
Full profile →

opencode / claude-opus-4-5

36.8%

95% interval 25.0–48.5 · 125 attempts

36.8% (95% interval 25.0–48.5)
Full profile →

Where they differ

Weaknessgpt-5.2claude-opus-4-5Difference
CWE-125Out-of-bounds Read70.3%41.9%+28.4
CWE-787Out-of-bounds Write75.0%41.2%+33.8
CWE-416Use After Free30.8%23.1%+7.7
CWE-908Use of Uninitialized Resource22.2%25.0%-2.8
CWE-415Double Free66.7%33.3%+33.4
CWE-1284Improper Validation of Specified Quantity in Input50.0%50.0%0.0
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0%0.0%+50.0
CWE-476NULL Pointer Dereference100.0%0.0%+100.0
CWE-4750.0%0.0%0.0
CWE-562100.0%0.0%+100.0

Per-class samples are small; a difference here is a hint about where to look, not a result. The corpus-level intervals above are the honest comparison.

Corpus cve-bench-136 · run February 2026 · How we test

See the benchmark →