Llama 4
Meta
partial detail 9 method fields published
Evaluation
Threat-modeling exercises identifying model capabilities necessary to automate operations or enhance human capabilities across key attack vectors, followed by developed challenges testing Llama 4 and peer models on automating cyberattacks, identifying/exploiting vulnerabilities, and automating harmful workflows
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 9 method fields published