Skip to content

Evaluation

Common use case and capability evaluations (system-level, with Llama Guard 3)

Common use case evaluations of systems composed of Llama models and Llama Guard 3 (input prompt/output response filtering) using dedicated adversarial evaluation datasets; capability evaluations with dedicated benchmarks for long context, multilingual, tool calls, coding, memorization

Harmful content & safety Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 3.1
Meta

partial detail 8 method fields published

See how this test was run
Llama 3.3
Meta

partial detail 8 method fields published

See how this test was run
Llama 4
Meta

partial detail 8 method fields published

See how this test was run