Skip to content

Evaluation

Internal AI research evaluation suite (7 environments)

ability to improve performance of ML code across LLMs, time series, low-level optimizations, reinforcement learning, and general problem solving

Autonomy & self-improvement Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 24 method fields published

See how this test was run