Skip to content

Evaluation

Refusal evaluations (Wildchat and XSTest)

refusal rates on toxic prompts (should refuse) and incorrect refusal rates on non-toxic prompts (should not refuse), using Wildchat and XSTest datasets

Refusal calibration Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku
Anthropic

partial detail 19 method fields published

See how this test was run
Claude 3.5 Sonnet
Anthropic

partial detail 15 method fields published

See how this test was run
Claude 3.5 Haiku, Claude 3.5 Sonnet (New)
Anthropic

partial detail 15 method fields published

See how this test was run