Evaluation
HumanEval
Ability to synthesize Python functions of varying complexity
Coding abilityCore capabilities
Anthropic, Google DeepMind — Gemini, OpenAI
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail
10 method fields published
See how this test was run
Gemini 1.0
Google DeepMind — Gemini
rich detail
8 method fields published
See how this test was run
Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku
Anthropic
partial detail
15 method fields published
See how this test was run
gpt-4o mini
OpenAI
rich detail
7 method fields published
See how this test was run