gpt-5.2-codex vs gpt-5.2
Both models ran the same corpus under the same verifier — codex and opencode respectively. The harness differs where the names differ, and that is part of what separates them.
Their confidence intervals overlap. On this corpus the difference of 2.4 points is not resolved — treat these two as tied.
Where they differ
| Weakness | gpt-5.2-codex | gpt-5.2 | Difference |
|---|---|---|---|
| CWE-125Out-of-bounds Read | 56.0% | 62.5% | -6.5 |
| CWE-787Out-of-bounds Write | 64.7% | 50.0% | +14.7 |
| CWE-416Use After Free | 15.4% | 33.3% | -17.9 |
| CWE-908Use of Uninitialized Resource | 22.2% | 22.2% | 0.0 |
| CWE-415Double Free | 33.3% | 66.7% | -33.4 |
| CWE-1284Improper Validation of Specified Quantity in Input | 50.0% | 0.0% | +50.0 |
| CWE-476NULL Pointer Dereference | 0.0% | 0.0% | 0.0 |
| CWE-843Access of Resource Using Incompatible Type ('Type Confusion') | 50.0% | 0.0% | +50.0 |
| CWE-475 | 100.0% | 0.0% | +100.0 |
| CWE-562 | 0.0% | 0.0% | 0.0 |
Per-class samples are small; a difference here is a hint about where to look, not a result. The corpus-level intervals above are the honest comparison.
Corpus cve-bench-136 · run February 2026 · How we test
See the benchmark →