Skip to main content

gpt-5-2-codex

run under the codex harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

49.2%

repair rate

36–60

95% interval

9

rank of 15

$6.26

per repair · estimated

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

63 repaired · 27 not repaired · 38 did not build

63 repaired · 27 not repaired · 38 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read56.0% (95% interval 41.0–66.3)56.0%75
CWE-787Out-of-bounds Write64.7% (95% interval 38.9–88.2)64.7%17
CWE-416Use After Free15.4% (95% interval 0.0–40.0)15.4%13
CWE-908Use of Uninitialized Resource22.2% — no interval: too few codebases or attempts to generalise22.2%9
CWE-415Double Free33.3% — no interval: too few codebases or attempts to generalise33.3%3
CWE-1284Improper Validation of Specified Quantity in Input50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-476NULL Pointer Dereference0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-475100.0% — no interval: too few codebases or attempts to generalise100.0%1
CWE-5620.0% — no interval: too few codebases or attempts to generalise0.0%1