Claude 3.5 Haiku, Claude 3.5 Sonnet (New)
Anthropic
partial detail 15 method fields published
Evaluation
ability to recognize adversarial prompts from users and behave in alignment with the system prompt (internal test sets of prompt injection attacks)
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 15 method fields published