Skip to content

Evaluation

Benchmark performance evolution during training

Tracking of model performance on a few QA and common sense benchmarks during training (Figure 2)

Core capabilities Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 1
Meta

partial detail 6 method fields published

See how this test was run