Research
Short notes on what we find when we read the results rather than rank them. Each one is computed from the published corpus and states what it does not show.
The samples nobody repaired
Pool every model and the corpus still does not close. The residue describes the frontier better than any single score.
A quarter of attempted repairs never compiled
Before a patch can be wrong it has to build. Which fraction fails depends more on the harness than the model.
The harness is part of the model
The same model scores points apart depending on the wrapper that runs it. A model-only ranking measures something it does not name.
Does AI help cyber attackers or defenders? (PDF)
A paired offense–defense differential on identical ground truth: the same vulnerabilities, asked to be triggered and repaired, scored by the same verifier.
Corpus cve-bench-136 · run February 2026 · See the benchmark