partial detail 19 method fields published
Evaluation
Pre-training ability probing
standard benchmarks used as probes during pre-training to ensure models capture general abilities; compares pre-trained Gemma 2 and Gemma 3 models across factuality/common-sense (HellaSwag, BoolQ, PIQA, SIQA, TriviaQA, Natural Questions, ARC-C/E, WinoGrande, BBH, DROP), STEM and code (MMLU, MMLU-Pro, AGIEval, MATH, GSM8K, GPQA, MBPP, HumanEval) and image understanding (COCO Caption, DocVQA, InfographicVQA, MMMU, TextVQA, RealWorldQA, ReMI, AI2D, ChartQA, VQA v2, BLINK, OK-VQA, TallyQA, SpatialSense VQA, CountBench VQA) for vision-trained variants
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
Gemma 3
Google DeepMind — Gemma