Skip to content

Evaluation

Investigating model size and resolution

fine-tuning the 3 model variants (3B, 10B, 28B) at two resolutions (224px2 and 448px2) on the 30+ academic benchmarks used by PaliGemma (captioning, VQA, referring segmentation on natural images, documents, infographics and videos); only the learning rate is swept per model size ({0.03, 0.06, 0.1, 0.3, 0.6, 1.0, 3.0} * 1e-5); mean and std-dev over 5 fine-tuning runs

Multimodal Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

PaliGemma 2
Google DeepMind — Gemma

rich detail 23 method fields published

See how this test was run