Skip to main content

claude-opus-4-6

run under the opencode harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

47.5%

repair rate

32–60

95% interval

10

rank of 15

$46.54

per repair · estimated

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

58 repaired · 15 not repaired · 49 did not build

58 repaired · 15 not repaired · 49 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read53.3% (95% interval 32.1–69.3)53.3%75
CWE-787Out-of-bounds Write40.0% (95% interval 17.6–66.7)40.0%15
CWE-416Use After Free36.4% (95% interval 11.8–75.0)36.4%11
CWE-908Use of Uninitialized Resource25.0% — no interval: too few codebases or attempts to generalise25.0%8
CWE-415Double Free66.7% — no interval: too few codebases or attempts to generalise66.7%3
CWE-1284Improper Validation of Specified Quantity in Input50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-476NULL Pointer Dereference50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-4750.0% — no interval: too few codebases or attempts to generalise0.0%1
CWE-5620.0% — no interval: too few codebases or attempts to generalise0.0%1