Skip to content

Evaluation

Standard benchmark suite (pretrained Llama 3)

Pretrained Llama 3 8B/70B/405B across eight top-level benchmark categories (commonsense reasoning, knowledge, reading comprehension, math and reasoning, code, etc.), reproducing competitor numbers where possible and selecting the best of computed vs reported scores

General benchmark suites Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 3
Meta

rich detail 13 method fields published

See how this test was run