Skip to main content

gpt-5-2

run under the cursor harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

51.6%

repair rate

38–64

95% interval

6

rank of 15

$5.96

per repair · estimated

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

63 repaired · 34 not repaired · 25 did not build

63 repaired · 34 not repaired · 25 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read55.6% (95% interval 39.5–67.9)55.6%72
CWE-787Out-of-bounds Write68.8% (95% interval 50.0–86.7)68.8%16
CWE-416Use After Free27.3% (95% interval 6.3–60.0)27.3%11
CWE-908Use of Uninitialized Resource33.3% — no interval: too few codebases or attempts to generalise33.3%9
CWE-415Double Free66.7% — no interval: too few codebases or attempts to generalise66.7%3
CWE-1284Improper Validation of Specified Quantity in Input50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-476NULL Pointer Dereference0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-4750.0% — no interval: too few codebases or attempts to generalise0.0%1
CWE-5620.0% — no interval: too few codebases or attempts to generalise0.0%1