Skip to main content

gpt-5.2 vs gemini-3.1-pro-preview

Both models ran the same corpus under the same verifier — codex and gemini31 respectively. The harness differs where the names differ, and that is part of what separates them.

Their confidence intervals overlap. On this corpus the difference of 4.0 points is not resolved — treat these two as tied.

codex / gpt-5.2

62.7%

95% interval 51.9–71.5 · 126 attempts

62.7% (95% interval 51.9–71.5)
Full profile →

gemini31 / gemini-3.1-pro-preview

58.7%

95% interval 44.0–71.2 · 109 attempts

58.7% (95% interval 44.0–71.2)
Full profile →

Where they differ

Weaknessgpt-5.2gemini-3.1-pro-previewDifference
CWE-125Out-of-bounds Read70.3%66.2%+4.1
CWE-787Out-of-bounds Write75.0%62.5%+12.5
CWE-416Use After Free30.8%33.3%-2.5
CWE-908Use of Uninitialized Resource22.2%0.0%+22.2
CWE-415Double Free66.7%100.0%-33.3
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0%50.0%0.0
CWE-476NULL Pointer Dereference100.0%50.0%+50.0
CWE-1284Improper Validation of Specified Quantity in Input50.0%100.0%-50.0
CWE-4750.0%100.0%-100.0
CWE-562100.0%0.0%+100.0

Per-class samples are small; a difference here is a hint about where to look, not a result. The corpus-level intervals above are the honest comparison.

Corpus cve-bench-136 · run February 2026 · How we test

See the benchmark →