Skip to content

Evaluation

Multi-lingual Benchmarks

code generation across a variety of popular programming languages measured with BabelCode (Orlanski et al. 2023) on BabelCode-translated HumanEval and MBPP datasets; languages include C++, C#, Go, Java, JavaScript, Kotlin, Python, Rust

Coding ability Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

CodeGemma
Google DeepMind — Gemma

partial detail 17 method fields published

See how this test was run