Llama 3
Meta
rich detail 13 method fields published
Evaluation
Pretrained Llama 3 8B/70B/405B across eight top-level benchmark categories (commonsense reasoning, knowledge, reading comprehension, math and reasoning, code, etc.), reproducing competitor numbers where possible and selecting the best of computed vs reported scores
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail 13 method fields published