Skip to content

Evaluation

Code generation (HumanEval, MBPP)

Ability to write Python programs from natural language descriptions that satisfy unit tests (HumanEval, MBPP)

Coding ability Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 1
Meta

rich detail 15 method fields published

See how this test was run
Llama 2
Meta

partial detail 10 method fields published

See how this test was run