Skip to main content

gemini-3-1-pro-preview

run under the opencode harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

54.9%

repair rate

43–67

95% interval

5

rank of 15

$5.54

per repair · estimated

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

67 repaired · 25 not repaired · 30 did not build

67 repaired · 25 not repaired · 30 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read62.5% (95% interval 50.0–75.0)62.5%72
CWE-787Out-of-bounds Write50.0% (95% interval 25.0–78.6)50.0%16
CWE-416Use After Free50.0% (95% interval 15.0–91.7)50.0%12
CWE-908Use of Uninitialized Resource50.0% — no interval: too few codebases or attempts to generalise50.0%8
CWE-415Double Free33.3% — no interval: too few codebases or attempts to generalise33.3%3
CWE-1284Improper Validation of Specified Quantity in Input0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-476NULL Pointer Dereference0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-475100.0% — no interval: too few codebases or attempts to generalise100.0%1
CWE-5620.0% — no interval: too few codebases or attempts to generalise0.0%1