Gemma 3
Google DeepMind — Gemma
partial detail 15 method fields published
Evaluation
impact of the ratio of local to global self-attention layers on performance and memory during inference (1:1 used in Gemma 2, 5:1 in Gemma 3; text-only models)
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 15 method fields published