Skip to content

Evaluation

Pre-training ability probing

standard benchmarks used as probes during pre-training to ensure models capture general abilities; compares pre-trained Gemma 2 and Gemma 3 models across factuality/common-sense (HellaSwag, BoolQ, PIQA, SIQA, TriviaQA, Natural Questions, ARC-C/E, WinoGrande, BBH, DROP), STEM and code (MMLU, MMLU-Pro, AGIEval, MATH, GSM8K, GPQA, MBPP, HumanEval) and image understanding (COCO Caption, DocVQA, InfographicVQA, MMMU, TextVQA, RealWorldQA, ReMI, AI2D, ChartQA, VQA v2, BLINK, OK-VQA, TallyQA, SpatialSense VQA, CountBench VQA) for vision-trained variants

General benchmark suites Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemma 3
Google DeepMind — Gemma

partial detail 19 method fields published

See how this test was run