claude-opus-4-6
run under the claude harness
The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.
61.6%
repair rate
49–71
95% interval
3
rank of 15
$2.69
per repair · metered
Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.
outcome mix
77 repaired · 28 not repaired · 20 did not build
By weakness class
| Weakness | Repair rate | Rate | Attempts |
|---|---|---|---|
| CWE-125Out-of-bounds Read | 69.3% | 75 | |
| CWE-787Out-of-bounds Write | 50.0% | 16 | |
| CWE-416Use After Free | 63.6% | 11 | |
| CWE-908Use of Uninitialized Resource | 11.1% | 9 | |
| CWE-415Double Free | 66.7% | 3 | |
| CWE-1284Improper Validation of Specified Quantity in Input | 50.0% | 2 | |
| CWE-843Access of Resource Using Incompatible Type ('Type Confusion') | 50.0% | 2 | |
| CWE-476NULL Pointer Dereference | 50.0% | 2 | |
| CWE-475 | 100.0% | 1 | |
| CWE-562 | 100.0% | 1 |
Compare
vs codex/gpt-5-2vs cursor/opus-4-6vs gemini31/gemini-3-1-pro-previewvs opencode/gemini-3-1-pro-preview
Corpus cve-bench-136 · run February 2026 · How we test