Skip to content

Evaluation

SWE-bench Verified

Human-validated subset of SWE-bench

Coding ability Anthropic, OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

o1-preview
OpenAI

rich detail 17 method fields published

See how this test was run
Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet
Anthropic

rich detail 19 method fields published

See how this test was run
o1
OpenAI

rich detail 17 method fields published

See how this test was run
o3-mini
OpenAI

rich detail 16 method fields published

See how this test was run
Claude 3.7 Sonnet
Anthropic

rich detail 29 method fields published

See how this test was run
Deep Research
OpenAI

rich detail 16 method fields published

See how this test was run
GPT-4.5
OpenAI

rich detail 15 method fields published

See how this test was run
GPT-4.1
OpenAI

rich detail 8 method fields published

See how this test was run
o3/o4-mini
OpenAI

rich detail 16 method fields published

See how this test was run
Claude Opus 4, Claude Sonnet 4
Anthropic

rich detail 30 method fields published

See how this test was run
GPT-5
OpenAI

rich detail 8 method fields published

See how this test was run