Llama 3.1
Meta
partial detail 8 method fields published
Evaluation
Common use case evaluations of systems composed of Llama models and Llama Guard 3 (input prompt/output response filtering) using dedicated adversarial evaluation datasets; capability evaluations with dedicated benchmarks for long context, multilingual, tool calls, coding, memorization
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 8 method fields published
partial detail 8 method fields published
partial detail 8 method fields published