Skip to content

Evaluation

HumanEval

Ability to synthesize Python functions of varying complexity

Coding abilityCore capabilities Anthropic, Google DeepMind — Gemini, OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

gpt-4
OpenAI

partial detail 10 method fields published

See how this test was run
Gemini 1.0
Google DeepMind — Gemini

rich detail 8 method fields published

See how this test was run
Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku
Anthropic

partial detail 15 method fields published

See how this test was run
gpt-4o mini
OpenAI

rich detail 7 method fields published

See how this test was run