Skip to content

Evaluation

Python Coding

performance on the canonical Python coding benchmarks HumanEval (Chen et al. 2021) and Mostly Basic Python Problems (Austin et al. 2021)

Coding ability Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

CodeGemma
Google DeepMind — Gemma

partial detail 22 method fields published

See how this test was run