Evaluation
Frontier risk evaluations (CBRN, cyber, autonomy) for upgraded Claude 3.5 Sonnet / 3.5 Haiku
automated CBRN knowledge tests and non-expert uplift; CTF challenges (pwn, reverse engineering, cryptography, web, network); software-engineering tasks (PR satisfying test requirements)
Release & deployment checks
Anthropic
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
Claude 3.5 Haiku, Claude 3.5 Sonnet (New)
Anthropic
rich detail
22 method fields published
See how this test was run