Out-of-bounds Write
The product writes data past the end, or before the beginning, of the intended buffer.
55.5%
repair rate
40–73
95% interval
247
scored attempts · 14 codebases
By model
| # | Agent | Model | Repair rate with interval | Rate | Attempts |
|---|---|---|---|---|---|
| 1 | cursor | opus-4.6 | 76.5% | 17 | |
| 2 | codex | gpt-5.2 | 75.0% | 16 | |
| 3 | cursor | gpt-5.2 | 68.8% | 16 | |
| 4 | codex | gpt-5.2-codex | 64.7% | 17 | |
| 5 | gemini31 | gemini-3.1-pro-preview | 62.5% | 16 | |
| 6 | cursor | composer-1.5 | 58.8% | 17 | |
| 7 | cursor | gpt-5.3-codex | 58.8% | 17 | |
| 8 | claude | claude-opus-4-6 | 50.0% | 16 | |
| 9 | opencode | gemini-3.1-pro-preview | 50.0% | 16 | |
| 10 | opencode | gpt-5.2 | 50.0% | 16 | |
| 11 | claude | claude-opus-4-5 | 47.1% | 17 | |
| 12 | gemini | gemini-3-pro-preview | 47.1% | 17 | |
| 13 | opencode | claude-opus-4-5 | 41.2% | 17 | |
| 14 | opencode | gpt-5.2-codex | 41.2% | 17 | |
| 15 | opencode | claude-opus-4-6 | 40.0% | 15 |
outcome mix
137 repaired · 52 not repaired · 58 did not build
Corpus cve-bench-136 · run February 2026 · How we test
All weakness classes →