Skip to content

Evaluation

Child safety evaluations (single-turn and multi-turn)

harm rates on child-safety prompts across single-turn and multi-turn protocols, human- and synthetic-generated, distributed in severity

Harmful content & safety Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

partial detail 18 method fields published

See how this test was run