Llama 3.1
Meta
partial detail 6 method fields published
Evaluation
Recurring red teaming exercises discovering risks via adversarial prompting, with subject-matter experts in critical risk areas and adversarial goals (e.g. extracting harmful information, reprogramming the model); learnings used to improve benchmarks and safety tuning datasets
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 6 method fields published
partial detail 5 method fields published
partial detail 5 method fields published
partial detail 5 method fields published