Llama 3
Meta
partial detail 12 method fields published
Evaluation
Final violation rate (VR) and false refusal rate (FRR) of Llama 3 405B vs similar models (two end-to-end API competitor systems and one internally hosted open-source model); the Llama 3 model card additionally notes internal false-refusal benchmarks and mitigations to make Llama 3 significantly less likely to falsely refuse than Llama 2
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 12 method fields published