Skip to content

Evaluation

Long context benchmark evaluation (RULER, MRCR)

performance of pre-trained and instruction fine-tuned models on long context benchmarks RULER and MRCR evaluated at 32K and 128K sequence lengths

Core capabilities Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemma 3
Google DeepMind — Gemma

partial detail 18 method fields published

See how this test was run