Skip to content

Evaluation

Trust & Safety model red-teaming (14 policy areas, 6 languages)

harm rates of responses across policy areas (Elections Integrity, Child Safety, Cyber Attacks, Hate & Discrimination, Violent Extremism, etc.) in English, Arabic, Spanish, Hindi, Tagalog, Chinese

Red teaming & adversarial testing Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.5 Haiku, Claude 3.5 Sonnet (New)
Anthropic

partial detail 18 method fields published

See how this test was run