Skip to main content

What a verified repair costs

A pass rate alone does not tell you what to run. The same corpus repaired by different models differs more in price than in capability, and no single model reaches most of it. These are the numbers a buyer actually decides on.

Cost per verified repair

Cost per successful repair spans $2.46 to $46.54 — a 18.9× spread. The most expensive model in this table is not the most capable one: price and capability are close to independent here.

Repair rate against cost per verified repair0%20%40%60%80%$5$10$20$50repair ratecost per verified repair (log scale)gpt-5.2: 62.7% at $4.91 per repair (estimated — bar shows the order-of-magnitude range)opus-4.6: 62.5% at $35.40 per repair (estimated — bar shows the order-of-magnitude range)claude-opus-4-6: 61.6% at $2.69 per repair (metered)gemini-3.1-pro-preview: 58.7% at $3.34 per repair (estimated — bar shows the order-of-magnitude range)gemini-3.1-pro-preview: 54.9% at $5.54 per repair (estimated — bar shows the order-of-magnitude range)gpt-5.2: 51.6% at $5.96 per repair (estimated — bar shows the order-of-magnitude range)gpt-5.2: 51.6% at $5.96 per repair (estimated — bar shows the order-of-magnitude range)gpt-5.3-codex: 50.4% at $6.11 per repair (estimated — bar shows the order-of-magnitude range)gpt-5.2-codex: 49.2% at $6.26 per repair (estimated — bar shows the order-of-magnitude range)claude-opus-4-6: 47.5% at $46.54 per repair (estimated — bar shows the order-of-magnitude range)claude-opus-4-5: 45.7% at $2.46 per repair (metered)composer-1.5: 45.2% at $3.87 per repair (estimated — bar shows the order-of-magnitude range)gemini-3-pro-preview: 43.0% at $4.56 per repair (estimated — bar shows the order-of-magnitude range)gpt-5.2-codex: 37.8% at $8.15 per repair (estimated — bar shows the order-of-magnitude range)claude-opus-4-5: 36.8% at $36.89 per repair (estimated — bar shows the order-of-magnitude range)
Filled points have metered token cost. Hollow points are estimated from turn counts and carry a bar showing the order-of-magnitude range that estimate supports — they are not precise to the cent. The dashed line joins models that nothing else beats on both price and rate.
AgentModelRepair rate95% interval$ / repairCost source
claudeclaude-opus-4-545.7%33.1–55.5$2.46metered
claudeclaude-opus-4-661.6%48.9–71.3$2.69metered
gemini31gemini-3.1-pro-preview58.7%44.0–71.2$3.34estimated
cursorcomposer-1.545.2%32.5–58.0$3.87estimated
geminigemini-3-pro-preview43.0%32.8–52.1$4.56estimated
codexgpt-5.262.7%51.9–71.5$4.91estimated
opencodegemini-3.1-pro-preview54.9%42.9–66.7$5.54estimated
cursorgpt-5.251.6%38.1–63.6$5.96estimated
opencodegpt-5.251.6%39.5–63.1$5.96estimated
cursorgpt-5.3-codex50.4%36.2–61.2$6.11estimated
codexgpt-5.2-codex49.2%35.9–59.8$6.26estimated
opencodegpt-5.2-codex37.8%25.0–48.7$8.15estimated
cursoropus-4.662.5%50.7–71.9$35.40estimated
opencodeclaude-opus-4-536.8%25.0–48.5$36.89estimated
opencodeclaude-opus-4-647.5%32.0–59.7$46.54estimated

Rates carry the same intervals as the board and overlap across most of this table — price separates these models far better than capability does. Only 2 of 15 models report metered token cost; the rest are estimated from turn counts and are marked as such. An estimate is not a measurement, so read the estimated rows as an order of magnitude, not a price.

How many models it takes to cover the corpus

Running more models only helps while they fail differently. Taking the models that add the most new repairs first, the best single model reaches 62.5% of the corpus; adding a second reaches 75%; the full set of 6 contributing models reaches 80.5%.

AddedModelNew repairsCumulative coverage
1cursor · opus-4.6+8062.5%
2codex · gpt-5.2+1675%
3codex · gpt-5.2-codex+377.3%
4opencode · gemini-3.1-pro-preview+278.9%
5claude · claude-opus-4-5+179.7%
6opencode · gpt-5.2-codex+180.5%

25 of 128 samples were never repaired by any model, in any attempt. That residue is the interesting part of this corpus: it is not a scoring artefact, it is where the whole field currently fails.

Publishing a number changes the game

Any published benchmark becomes something to optimise against, and a score that has been trained on stops measuring capability. We keep a permanent held-out slice that never appears in any published figure on this site, and we say so here rather than claiming the published set is the whole corpus.

The same reasoning applies to which side of security AI helps. Measuring repair alone answers half the question; the other half is whether the same systems are better at breaking than fixing. Our paired study measured both on identical ground truth and found the balance favoured the attacker side by roughly eleven points on average, with the direction depending on the system. Read the paper (PDF).

Corpus cve-bench-136 · run February 2026 · How we test