Skip to content

Evaluation

Scaled evaluations with Purple Llama safeguards

Scaled evaluations using dedicated adversarial evaluation datasets, evaluating systems composed of Llama models and Purple Llama safeguards for input prompt and output response filtering

Harmful content & safety Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 3.2
Meta

partial detail 8 method fields published

See how this test was run