Skip to content

Evaluation

HumanEval Infilling

single-line and multi-line metrics in the HumanEval Infilling benchmarks introduced in Fried et al. (2023); latency measured as total seconds to obtain 128-token continuations per task

Coding ability Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

CodeGemma
Google DeepMind — Gemma

rich detail 26 method fields published

See how this test was run