Skip to main content

claude-opus-4-6 vs composer-1.5

Both models ran the same corpus under the same verifier — claude and cursor respectively. The harness differs where the names differ, and that is part of what separates them.

Their confidence intervals overlap. On this corpus the difference of 16.4 points is not resolved — treat these two as tied.

claude / claude-opus-4-6

61.6%

95% interval 48.9–71.3 · 125 attempts

61.6% (95% interval 48.9–71.3)
Full profile →

cursor / composer-1.5

45.2%

95% interval 32.5–58.0 · 126 attempts

45.2% (95% interval 32.5–58.0)
Full profile →

Where they differ

Weaknessclaude-opus-4-6composer-1.5Difference
CWE-125Out-of-bounds Read69.3%53.4%+15.9
CWE-787Out-of-bounds Write50.0%58.8%-8.8
CWE-416Use After Free63.6%15.4%+48.2
CWE-908Use of Uninitialized Resource11.1%11.1%0.0
CWE-415Double Free66.7%33.3%+33.4
CWE-1284Improper Validation of Specified Quantity in Input50.0%50.0%0.0
CWE-843Access of Resource Using Incompatible Type ('Type Confusion')50.0%50.0%0.0
CWE-476NULL Pointer Dereference50.0%0.0%+50.0
CWE-475100.0%0.0%+100.0
CWE-562100.0%0.0%+100.0

Per-class samples are small; a difference here is a hint about where to look, not a result. The corpus-level intervals above are the honest comparison.

Corpus cve-bench-136 · run February 2026 · How we test

See the benchmark →