Skip to content

Evaluation

Appropriate harmlessness evaluation

four-way categorization of responses (helpful answer, policy violation, appropriate refusal, unnecessary refusal) using refusal and policy-violation classifiers plus maximally-helpful reference responses

Refusal calibration Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 20 method fields published

See how this test was run