Skip to main content

gemini-3-pro-preview

run under the gemini harness

The harness is part of what is measured here: the same model under a different runner scores differently. Why that matters.

43.0%

repair rate

33–52

95% interval

13

rank of 15

$4.56

per repair · estimated

Rank is a position in a table whose intervals overlap almost everywhere. Read it as a band, not a place.

outcome mix

55 repaired · 36 not repaired · 37 did not build

55 repaired · 36 not repaired · 37 did not build

By weakness class

WeaknessRepair rateRateAttempts
CWE-125Out-of-bounds Read53.3% (95% interval 39.7–63.3)53.3%75
CWE-787Out-of-bounds Write47.1% (95% interval 25.0–68.8)47.1%17
CWE-416Use After Free23.1% (95% interval 6.7–44.4)23.1%13
CWE-908Use of Uninitialized Resource11.1% — no interval: too few codebases or attempts to generalise11.1%9
CWE-415Double Free33.3% — no interval: too few codebases or attempts to generalise33.3%3
CWE-1284Improper Validation of Specified Quantity in Input50.0% — no interval: too few codebases or attempts to generalise50.0%2
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-476NULL Pointer Dereference0.0% — no interval: too few codebases or attempts to generalise0.0%2
CWE-4750.0% — no interval: too few codebases or attempts to generalise0.0%1
CWE-5620.0% — no interval: too few codebases or attempts to generalise0.0%1