Skip to content

Evaluation

SWE-bench Verified (hard subset)

ability to resolve real-world GitHub issues (42 hard tasks estimated to require >1 hour of engineering work) as a precursor to autonomy

Coding ability Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 26 method fields published

See how this test was run
Claude Opus 4, Claude Sonnet 4
Anthropic

rich detail 25 method fields published

See how this test was run