Skip to content

Evaluation

MMLU

Massive Multitask Language Understanding: 57 subjects of multiple-choice problems

Core capabilities Anthropic, Meta, OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 1
Meta

partial detail 11 method fields published

See how this test was run
gpt-4
OpenAI

rich detail 14 method fields published

See how this test was run
Claude 2
Anthropic

rich detail 20 method fields published

See how this test was run
Llama 2
Meta

partial detail 10 method fields published

See how this test was run
Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku
Anthropic

rich detail 16 method fields published

See how this test was run
gpt-4o mini
OpenAI

partial detail 7 method fields published

See how this test was run