Skip to main content

claude-opus-4-6

run under the claude harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

61.6%

repair rate

49–71

95% interval

3

rank of 15

$2.69

per repair · metered

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

77 repaired · 28 not repaired · 20 did not build

77 repaired · 28 not repaired · 20 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read69.3% (95% interval 52.9–81.0)69.3%75
CWE-787Out-of-bounds Write50.0% (95% interval 27.8–73.3)50.0%16
CWE-416Use After Free63.6% (95% interval 25.0–100.0)63.6%11
CWE-908Use of Uninitialized Resource11.1% — no interval: too few codebases or attempts to generalise11.1%9
CWE-415Double Free66.7% — no interval: too few codebases or attempts to generalise66.7%3
CWE-1284Improper Validation of Specified Quantity in Input50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-476NULL Pointer Dereference50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-475100.0% — no interval: too few codebases or attempts to generalise100.0%1
CWE-562100.0% — no interval: too few codebases or attempts to generalise100.0%1