Skip to main content

gpt-5-2

run under the codex harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

62.7%

repair rate

52–72

95% interval

1

rank of 15

$4.91

per repair · estimated

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

79 repaired · 12 not repaired · 35 did not build

79 repaired · 12 not repaired · 35 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read70.3% (95% interval 57.1–79.2)70.3%74
CWE-787Out-of-bounds Write75.0% (95% interval 52.6–100.0)75.0%16
CWE-416Use After Free30.8% (95% interval 12.5–50.0)30.8%13
CWE-908Use of Uninitialized Resource22.2% — no interval: too few codebases or attempts to generalise22.2%9
CWE-415Double Free66.7% — no interval: too few codebases or attempts to generalise66.7%3
CWE-1284Improper Validation of Specified Quantity in Input50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-476NULL Pointer Dereference100.0% — no interval: too few codebases or attempts to generalise100.0%2
CWE-4750.0% — no interval: too few codebases or attempts to generalise0.0%1
CWE-562100.0% — no interval: too few codebases or attempts to generalise100.0%1