Evaluation
SWE-bench Verified (hard subset)
ability to resolve real-world GitHub issues (42 hard tasks estimated to require >1 hour of engineering work) as a precursor to autonomy
Coding ability
Anthropic
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
Claude 3.7 Sonnet
Anthropic
rich detail
26 method fields published
See how this test was run
Claude Opus 4, Claude Sonnet 4
Anthropic
rich detail
25 method fields published
See how this test was run