Evaluation
SWE-bench Verified
Human-validated subset of SWE-bench
Coding ability
Anthropic, OpenAI
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
o1-preview
OpenAI
rich detail
17 method fields published
See how this test was run
Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet
Anthropic
rich detail
19 method fields published
See how this test was run
rich detail
17 method fields published
See how this test was run
rich detail
16 method fields published
See how this test was run
Claude 3.7 Sonnet
Anthropic
rich detail
29 method fields published
See how this test was run
Deep Research
OpenAI
rich detail
16 method fields published
See how this test was run
rich detail
15 method fields published
See how this test was run
rich detail
8 method fields published
See how this test was run
o3/o4-mini
OpenAI
rich detail
16 method fields published
See how this test was run
Claude Opus 4, Claude Sonnet 4
Anthropic
rich detail
30 method fields published
See how this test was run
rich detail
8 method fields published
See how this test was run