Claude 3.7 Sonnet
Anthropic
rich detail 26 method fields published
Evaluation
hard subset of SWE-bench Verified (2-8 hour software engineering tasks) and custom difficult AI R&D tasks built in-house, with reference expert solutions and difficulty variants
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail 26 method fields published