Skip to content

Evaluation

Autonomy evaluations (Claude 3.7 Sonnet RSP)

hard subset of SWE-bench Verified (2-8 hour software engineering tasks) and custom difficult AI R&D tasks built in-house, with reference expert solutions and difficulty variants

Autonomy & self-improvement Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 26 method fields published

See how this test was run