← Back to Explorer
Evaluation 567 of 723
Evaluation
Refusal evaluations (Wildchat and XSTest)
refusal rates on toxic prompts (should refuse) and incorrect refusal rates on non-toxic prompts (should not refuse), using Wildchat and XSTest datasets
Refusal calibration
Anthropic
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku
Mar 2024
Anthropic
partial detail
19 method fields published
See how this test was run
The Claude 3 Model Family: Opus, Sonnet, Haiku · Mar 2024
The Claude 3 Model Family: Opus, Sonnet, Haiku · Mar 2024
The Claude 3 Model Family: Opus, Sonnet, Haiku · Mar 2024
Claude 3.5 Sonnet
Jun 2024
Anthropic
partial detail
15 method fields published
See how this test was run
The Claude 3 Model Family: Opus, Sonnet, Haiku · Mar 2024
The Claude 3 Model Family: Opus, Sonnet, Haiku · Mar 2024
Claude 3.5 Sonnet Model Card Addendum | Anthropic · Jun 2024
Claude 3.5 Haiku, Claude 3.5 Sonnet (New)
Oct 2024
Anthropic
partial detail
15 method fields published
See how this test was run
Claude 3.5 Sonnet Model Card Addendum | Anthropic · Jun 2024
Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet · Oct 2024
Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet · Oct 2024