Skip to content

Evaluation

Frontier risk evaluations (CBRN, cyber, autonomy) for Claude 3.5 Sonnet

per-domain quantitative thresholds of concern across CBRN, cybersecurity, and autonomous capabilities; refusal rates measured and non-refusal elicitation used to estimate helpful-only performance

Release & deployment checks Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.5 Sonnet
Anthropic

rich detail 24 method fields published

See how this test was run