Llama 3
Meta
partial detail 13 method fields published
Evaluation
Violation and false refusal rates on DocQA and Many-shot long-context benchmarks; mitigation via SFT datasets with safe behavior amid in-context unsafe demonstrations and a scalable mitigation strategy
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 13 method fields published